Distributed training — DDP, ZeRO, FSDP, and what actually crosses the wire¶
Every model in this book eventually outgrows one GPU — either its compute is too slow for the deadline or its parameters no longer fit in one card's memory — and the fix is always some flavor of putting more silicon on the problem. More silicon means more silicon that has to talk to itself, and that conversation is not free: it crosses real wires at a real, finite bandwidth, and it can dominate a training step's wall-clock time as easily as it can vanish into a well-timed overlap. …
🔒 La suite est en accès freemium — lecture complète 100 % gratuite
Tu lis ici l'aperçu libre. Le reste du chapitre (code, schémas, maths, exercices) fait partie du livre complet : crée un compte gratuit (30 secondes, aucun paiement) pour tout lire.