Skip to content

Serving LLMs at scale — KV cache, continuous batching, and dollars per million tokens

Every chapter before this one produced a trained model and stopped at the boundary where a request meets it — Chapter 16 named batching, quantization, and distillation as generic serving levers without pricing an LLM specifically, and Chapter 19 derived the KV cache and sketched continuous batching and PagedAttention in prose, in the middle of a chapter about training and evaluating the model itself. This chapter is the one that answers what a "deploy and monitor LLMs in production" job posting actually means: given a model, a GPU, and a latency SLO, how many concurrent users fit in memory, how many tokens per second you can sell, and what each million of them costs. …

🔒 La suite est en accès freemium — lecture complète 100 % gratuite

Tu lis ici l'aperçu libre. Le reste du chapitre (code, schémas, maths, exercices) fait partie du livre complet : crée un compte gratuit (30 secondes, aucun paiement) pour tout lire.

Créer un compte gratuit Se connecter