Book Podcast Courses Playlists Writing Speaking Wall of Love About Mani Research AWS Blogs

Subscribe

LinkedIn YouTube Instagram Substack TikTok GitHub

Free course

Inference Optimization

Training is a one time capital cost. Inference is the bill that arrives every month. This course is about the engineering that moves it.

  • Free on YouTube
  • Five to ten minute chapters
Mani Khanuja explaining a concept from her studio

What this is

Latency is a product decision.

Training is a one time capital cost. Inference is the recurring one, and it is where the engineering leverage now sits.

The spine of the whole course is one number: arithmetic intensity. Every technique either raises it, lowers the bytes it divides by, or hides latency behind it.

Episodes run five to ten minutes. Every module closes with a CLINIC episode, where the symptom comes first and the theory is used to find the cause.

Who it is for

  • Beginners starting out
  • Curious engineers
  • People interviewing for inference and ML systems roles
  • Professionals reviewing or deep diving
Open the playlist

Why it matters

One slide justifies the whole course.

One slide justifies the entire course. DeepSeek published 24 hours of its own production economics:

608B
input tokens served in a day
168B
output tokens served in a day
226.75
H800 nodes, on average
$87,072
cost per day, against $562,027 of theoretical revenue

The shape of it

Every technique is one of four levers.

Whatever the trick is called, it does one of these four things. Once you can name which, you can reason about whether it will help you.

Lever 01

Move fewer bytes

Cut the denominator of arithmetic intensity.

Lever 02

Do fewer FLOPs

Cut real work.

Lever 03

Share work across requests

Amortise a fixed cost.

Lever 04

Overlap

Hide one resource behind another.

Finished a chapter? Go deeper.

The newsletter goes further on the same material, the agents course builds on it, and the podcast argues the business side.