Batch processing — Spark internals, dbt, DuckDB, idempotence¶
Most of the world's data work is still batch: a bounded input, processed on a schedule, producing a bounded output. Streaming gets the conference talks, but the nightly job that rebuilds your warehouse, the hourly job that compacts your lake, and the backfill that reprocesses six months of history after a bug — those are batch, and they move orders of magnitude more bytes than your streaming topology does. …
🔒 La suite est en accès freemium — lecture complète 100 % gratuite
Tu lis ici l'aperçu libre. Le reste du chapitre (code, schémas, maths, exercices) fait partie du livre complet : crée un compte gratuit (30 secondes, aucun paiement) pour tout lire.