Skip to content

Glossary (exhaustive)

This appendix is the single alphabetical reference for every term used across the book, from cgroups to the Bellman equation. Each entry does three things: it defines the term precisely, it explains why it matters in the workflow the book teaches, and — whenever the term has a mathematical basis — it derives the formula and computes it on real numbers, so you never have to take a symbol on faith. Full multi-step derivations (SVD, backpropagation, KL divergence family, Bellman optimality) live in Appendix D; this glossary gives you the compressed, load-bearing version you can recall in a design review or an interview.

Mental model. The book's arc is DevOps → MLOps → AI: you first learn to run any workload reliably (Linux, networking, git, containers, Kubernetes, IaC, CI/CD, security, observability — the platform layer); then you learn to run data and model workloads reliably on that platform (Python data stack, data engineering, DVC, MLflow, MLOps patterns — the ML platform layer); then you learn the models themselves (classical ML, deep learning, NLP/LLMs, computer vision, generative AI, reinforcement learning, time series — the applied AI layer). Every term below belongs to exactly one of these layers, and most of the confusion between "the ML person" and "the platform person" on a team comes from silently assuming the other one means the same thing by a shared word (state, policy, pipeline, deployment). After this appendix you will be able to: (1) place any unfamiliar term into the right layer in one glance, (2) reconstruct the formula behind it without looking it up, and (3) spot when two team members are using the same word for two different concepts.

flowchart TD
    subgraph Platform["Platform layer — DevOps"]
        L1["Linux & Shell (Ch.1)"] --> L2["Networking (Ch.2)"]
        L2 --> L3["Git (Ch.3)"]
        L3 --> L4["Docker (Ch.4)"]
        L4 --> L5["AWS (Ch.5)"]
        L5 --> L6["Kubernetes (Ch.6)"]
        L6 --> L7["Terraform (Ch.7)"]
        L7 --> L8["CI/CD (Ch.8)"]
        L8 --> L9["Security (Ch.9)"]
        L9 --> L10["Observability (Ch.10)"]
    end
    subgraph MLPlatform["ML platform layer — MLOps"]
        L11["Python data stack (Ch.11)"] --> L12["Data engineering (Ch.12)"]
        L12 --> L13["ML foundations (Ch.13)"]
        L13 --> L14["DVC (Ch.14)"]
        L14 --> L15["MLflow (Ch.15)"]
        L15 --> L16["MLOps patterns (Ch.16)"]
    end
    subgraph AppliedAI["Applied AI layer"]
        L17["Classical ML (Ch.17)"]
        L18["Deep learning (Ch.18)"]
        L19["NLP & LLMs (Ch.19)"]
        L20["Computer vision (Ch.20)"]
        L21["Generative AI (Ch.21)"]
        L22["Reinforcement learning (Ch.22)"]
        L23["Time series (Ch.23)"]
    end
    Platform --> MLPlatform --> AppliedAI --> Capstone["Capstone (Ch.24)"]

How to use this glossary

Entries are grouped into four alphabetical bands (A–C, D–H, I–O, P–Z) purely to keep each band scrollable; within a band, terms are strictly alphabetical. Each entry ends with the chapter(s) that develop it in full — click through when you need the complete treatment, code, and exercises. A recurring pattern you will see below is define → derive → compute: take Little's law as the smallest possible illustration. It states that in steady state, the average number of items in a system equals the arrival rate times the average time an item spends in the system:

\[L = \lambda W\]

Here \(L\) is the average number of concurrent requests "in flight," \(\lambda\) is the arrival rate (requests per second), and \(W\) is the average time each request spends in the system (seconds). If your API receives \(\lambda = 50\) requests/s and each takes on average \(W = 0.2\,\text{s}\) to answer, then on average \(L = 50 \times 0.2 = 10\) requests are being processed concurrently at any instant — which is exactly the number of worker threads/pods you must be able to run in parallel to avoid queueing. Every math-bearing entry below follows this same three-step discipline.

A–C

  • ACID — Atomicity, Consistency, Isolation, Durability: the four guarantees a transactional data store makes so that concurrent writes never leave the database in a half-updated, contradictory state. Atomicity means a transaction is all-or-nothing; Isolation means concurrent transactions don't see each other's half-finished work; Durability means a committed write survives a crash. Lakehouse formats (Delta Lake, Iceberg, Hudi) exist precisely to bring ACID semantics to files sitting on object storage, which historically had none. (Ch. 12)
  • Advantage \(A(s,a)\) — how much better than the average action a specific action \(a\) is in state \(s\): \(A(s,a) = Q(s,a) - V(s)\), where \(Q(s,a)\) is the expected return of taking \(a\) in \(s\) and \(V(s)\) is the expected return of state \(s\) under the current policy. Worked example: if \(Q(s,a)=10\) and \(V(s)=7\), then \(A(s,a)=3\): this action beats the state's average outcome by 3 units of return, so a policy-gradient update should increase its probability. Using the advantage instead of the raw return \(Q(s,a)\) subtracts a state-dependent baseline, which does not bias the gradient's expectation but sharply reduces its variance — the single most important variance-reduction trick in policy-gradient methods. (Ch. 22)
  • Agent loop (state machine) — the ReAct-style PLAN → ACT → OBSERVE → (PLAN | DONE | FAILED) cycle modeled as a finite state machine with named states and guarded transitions, instead of an unstructured while loop whose only implicit state is "still running." DONE and FAILED are absorbing: no transition leaves them, and every arrow into them is a checked terminal condition, not an inferred one. (Ch. 19e)
  • AKS (Azure Kubernetes Service) — Azure's managed Kubernetes offering: Microsoft runs the control plane at no charge (versus EKS's per-cluster hourly fee), you still pay for the node VMs underneath it. Ships the same Kubernetes object model Chapter 6 builds against, with the provisioning story swapped: az aks create / the azurerm_kubernetes_cluster Terraform resource in place of eksctl / aws_eks_cluster. (Ch. 16g)
  • AMI (Amazon Machine Image) — the frozen template (root filesystem + boot config) an EC2 instance is launched from. Bake application dependencies into a custom AMI (via Packer, for example) to cut cold-boot time versus provisioning everything with a startup script every time. (Ch. 5)
  • AMP / Autocast (Automatic Mixed Precision) — the mechanism that runs matmuls/convolutions in fp16 or bf16 (fast on tensor cores) while keeping a small set of numerically sensitive ops (reductions, softmax, loss) in fp32, inside a torch.autocast / tf.keras.mixed_precision context. AMP does not by itself halve weight-and-optimizer memory — that stays at 16 bytes/parameter for Adam either way — its win is compute throughput and activation memory. With fp16, AMP is paired with a GradScaler (loss scaling) to keep small gradients from underflowing fp16's narrow exponent range; bf16, with the same 8 exponent bits as fp32, needs no loss scaling at all. (Ch. 16c)
  • ANN (Approximate Nearest Neighbor) search / vector database — the retrieval engine behind RAG (below): embeddings (below) are high-dimensional dense vectors, and finding the true \(k\) nearest neighbors of a query vector by brute force is \(O(n)\) per query over \(n\) stored vectors — far too slow once \(n\) is in the millions. ANN indexes (HNSW — Hierarchical Navigable Small World graphs, or IVF — Inverted File index with product quantization, as used by FAISS) trade a small, tunable amount of recall for query time close to \(O(\log n)\) by pre-building a graph or clustering structure that lets search skip most of the vector space. A vector database (Pinecone, Weaviate, pgvector, Qdrant) packages an ANN index with metadata filtering, persistence, and horizontal scaling into a queryable service. Similarity is usually cosine similarity, \(\cos(\theta) = \dfrac{u\cdot v}{\lVert u\rVert\lVert v\rVert}\): worked example, \(u=[1,2]\), \(v=[2,1]\) give \(u\cdot v = 1\times2+2\times1=4\), \(\lVert u\rVert=\lVert v\rVert=\sqrt5\approx2.236\), so \(\cos(\theta)=4/(2.236\times2.236)=4/5=0.8\) — fairly similar, but not identical, direction. (Ch. 19)
  • Ansible — an agentless, push-model configuration-management tool: the control node opens an SSH connection to each target, copies over a small Python payload, executes it, and tears it down again — no long-running daemon installed on the target, unlike Chef or Puppet's pull agents. A playbook (below) declares the desired state of installed packages, files, and services; each module it calls implements a get-then-set contract (check current state, only change what differs), which is what makes a playbook idempotent (see Idempotency, above) — the same run against an already-converged host reports zero changes. Terraform (Ch. 7) provisions the machine; Ansible takes over exactly where terraform apply stops, configuring what runs on it. (Ch. 16f)
  • Ansible Vault — Ansible's built-in mechanism for encrypting a value or an entire file at rest, inside the playbooks and inventories a git repository stores (AES256, keyed by a vault password or password-derivation script). It solves a narrower problem than a runtime secrets manager: Vault protects a secret sitting in version control, but rotating it still requires a human to re-encrypt the file and re-run the playbook — it has no refreshInterval of its own. ESO (above) and its Kubernetes refreshInterval protect a secret an application reads at runtime; conflating the two boundaries means a "rotated" credential is only ever as fresh as the last person who remembered to redeploy. (Ch. 16f)
  • API design (REST vs. gRPC vs. GraphQL) — three common contracts for a service boundary. REST models a system as resources addressed by URLs and manipulated with HTTP verbs (GET/POST/PUT/DELETE), is human-readable (usually JSON) and cacheable by ordinary HTTP infrastructure, but tends to either over-fetch (an endpoint returns fields the client doesn't need) or under-fetch (the client must chain several calls). GraphQL flips this: the client sends a query describing exactly the fields/relations it wants, and the server resolves precisely that shape in one round trip — solving over/under-fetching at the cost of losing simple HTTP caching and needing query-complexity limits to prevent a single request from fanning out into an expensive resolver tree. gRPC compiles a strongly-typed schema (Protocol Buffers) into client/server stubs in every target language, communicates over HTTP/2 with binary serialization (much smaller and faster to parse than JSON), and natively supports bidirectional streaming — the default choice for low-latency internal service-to-service calls, at the cost of not being directly browser- or curl-friendly. (Ch. 2)
  • ARIMA(p, d, q) — AutoRegressive Integrated Moving Average, the classical linear model for a stationary (after differencing \(d\) times) time series: it regresses the current value on its own past \(p\) values (AR term) and on past forecast errors' \(q\) terms (MA term). It remains a strong, cheap, interpretable baseline that every deep-learning forecaster must beat before being trusted in production. (Ch. 23)
  • ARN (Amazon Resource Name) — the globally unique identifier of any AWS resource, e.g. arn:aws:s3:::my-bucket/key. IAM policies grant or deny actions on ARNs (with wildcards), never on human-friendly names, which is why a typo'd ARN in a policy silently grants access to nothing rather than erroring. (Ch. 5)
  • ASGI (vs WSGI) — the two Python web-server interface standards. WSGI (Web Server Gateway Interface) is synchronous: the server calls one callable per request and that callable blocks the calling thread until it returns, so concurrency comes only from running more OS threads/processes. ASGI (Asynchronous Server Gateway Interface — what FastAPI/Starlette/uvicorn implement) instead calls an async def coroutine that can suspend at an await and let the event loop (below) run other pending work in the meantime, on a single thread — the interface a multi-second AI request needs, since holding hundreds of slow requests as coroutines costs kilobytes each, versus megabytes of reserved stack per OS thread under WSGI. (Ch. 19a)
  • Attack success rate (ASR) — the fraction of a labeled adversarial prompt set that achieves its attacker-defined goal against a given defense configuration, measured by running the same fixed corpus through a pipeline before and after each guardrail layer is added. It is the empirical check on a defense-in-depth independence assumption: the product-of-miss-rates arithmetic predicts what a layered stack's pass-through should be if the layers are independent, and ASR tells you what it actually is. (Ch. 19f)
  • Attention — a content-based routing mechanism that lets each output position look at every input position and weigh them by relevance, rather than only at a fixed nearby window. Scaled dot-product attention is $\(\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V,\)$ where \(Q\) (queries), \(K\) (keys), \(V\) (values) are learned linear projections of the input, and \(\sqrt{d_k}\) rescales the dot products so softmax doesn't saturate for large key dimension \(d_k\). Worked toy example: with \(d_k=2\), query \(Q=[1,0]\) and keys \(K_1=[1,0]\), \(K_2=[0,1]\), the raw scores are \(Q\cdot K_1/\sqrt2 = 1/1.414 = 0.707\) and \(Q\cdot K_2/\sqrt2 = 0\). Softmax over \([0.707, 0]\) gives \(e^{0.707}=2.028\) and \(e^{0}=1\), sum \(=3.028\), so the attention weights are \([0.670,\,0.330]\): token 1 receives 67% of the attention, token 2 the remaining 33%. This is the core primitive of the Transformer. (Ch. 19)
  • AUC / ROC curve — the Receiver Operating Characteristic plots the true-positive rate against the false-positive rate as a classifier's decision threshold sweeps from 0 to 1; the Area Under that Curve is a single threshold-independent number between 0.5 (random) and 1.0 (perfect ranking). AUC equals the probability that a randomly chosen positive example is scored higher than a randomly chosen negative one — useful when the operating threshold isn't fixed yet, but it can look deceptively good on very imbalanced data (prefer precision-recall AUC there). (Ch. 13, Ch. 17)
  • Autoencoder — a network trained to reconstruct its own input through a narrow bottleneck, forcing it to learn a compressed representation. The plain (deterministic) autoencoder learns compression only; the Variational Autoencoder (VAE, below) turns the bottleneck into a probability distribution so you can sample new, never-seen data from it. (Ch. 18, Ch. 21)
  • Autoscaling (HPA / VPA / CA) — the family of Kubernetes controllers that add/remove replicas (Horizontal Pod Autoscaler), resize a pod's CPU/memory requests (Vertical Pod Autoscaler), or add/remove worker nodes (Cluster Autoscaler) in response to observed load. HPA and VPA should not target the same metric on the same workload — they will fight each other. (Ch. 6)
  • Autovacuum threshold — the dead-tuple count that triggers Postgres's background vacuum on a table, \(\text{trigger} = \texttt{autovacuum\_vacuum\_threshold} + \texttt{autovacuum\_vacuum\_scale\_factor} \times \text{reltuples}\) (defaults 50 and 0.2). Both GUCs are overridable per table — the mechanism a high-churn ingestion table needs, because the cluster default (tuned for an average table) waits far too long: at 300,000 rows the default trigger is 60,050 dead tuples, versus 6,500 at a tuned 2% scale factor, roughly 9× more frequent passes each doing far less work. (Ch. 19b)
  • AWQ / GPTQ — two post-training weight-quantization schemes for LLM inference, both typically targeting INT4. GPTQ quantizes column-by-column, correcting the remaining columns after each one to compensate for the rounding error just introduced (a second-order, Hessian-informed correction). AWQ instead leaves a small fraction of the most activation-sensitive weight channels at higher precision (or scales them before quantizing) and quantizes the rest uniformly, which costs almost nothing in size but protects exactly the channels that dominate output error under real activation distributions. Both are calibration-set-dependent: a checkpoint quantized against general web text can regress on a domain-shifted deployment its calibration set never saw. (Ch. 16e, Ch. 19)
  • AZ (Availability Zone) — one or more physically isolated datacenters within an AWS Region, each with independent power/cooling/networking. Spreading instances across \(\ge 2\) AZs is the cheapest form of high availability: an AZ-wide outage (power, fiber cut) takes down at most half your fleet. (Ch. 5)
  • Backfill scheduling (Slurm) — a scheduling pass Slurm's slurmctld runs alongside strict priority-order scheduling: for every idle-but-reserved resource, it looks for a pending, lower-priority job small or short enough to run to completion in the gap before a higher-priority job's reserved start time, without delaying that reservation by even one second. Worked example: a node frees fully in 90 minutes for a reserved 8-GPU job; a pending 2-GPU job declares --time=01:00:00 — \(60 < 90\), so it runs now and vacates with 30 minutes to spare. This is why a smaller, lower-priority job can legitimately start before a larger, higher-priority one still waiting in the queue — it costs the larger job nothing. (Ch. 16b)
  • Backoff / retry / jitter — the standard pattern for calling a flaky or overloaded remote dependency: on failure, retry, but wait progressively longer between attempts (exponential backoff, \(\text{wait}_n = \min(\text{cap},\, \text{base}\times 2^n)\)) so a struggling downstream service isn't hit harder the moment it starts failing. Naive exponential backoff still synchronizes many clients' retries into the same instant (they all failed at once and all back off by the same schedule), which can produce a "retry storm" wave — jitter breaks this by randomizing the wait within a range, e.g. \(\text{wait}_n = \text{random}(0,\, \text{base}\times 2^n)\) ("full jitter"). Worked example: with base \(=100\,\text{ms}\) and cap \(=10\,\text{s}\), attempt \(n=5\) gives an un-jittered wait of \(\min(10000, 100\times2^5)=\min(10000,3200)=3200\,\text{ms}\); full jitter instead draws uniformly from \([0, 3200]\,\text{ms}\), spreading a thundering herd of simultaneous retries across more than three seconds instead of concentrating them at exactly 3.2 s. Always pair retries with a maximum attempt count and, ideally, a circuit breaker (below) so a permanently down dependency doesn't retry forever. (Ch. 2)
  • Backpressure (bounded queue + load shedding) — an explicit limit on how much unprocessed work a service will admit, versus letting an unbounded queue absorb every burst silently. A bounded asyncio.Queue(maxsize=N) that returns an immediate error (HTTP 503) once full converts an invisible, ever-growing wait into a fast, actionable failure the caller can retry elsewhere; the alternative — an unbounded queue, or a bounded queue whose producer blocks on put() instead of failing on put_nowait() — turns overload into a p99 latency graph that creeps upward with no error and no exception to alert on. Worked example: a queue capped at \(Q_\text{max}=100\) draining at \(C=150\) req/s under a burst of \(\lambda_\text{burst}=300\) req/s fills at a net \(300-150=150\) req/s, giving a worst-case wait of \(100/150\approx0.67\,\text{s}\) for any accepted request — bounded and predictable, unlike an unbounded queue's unbounded wait. (Ch. 19a)
  • Backpropagation — the algorithm that computes \(\partial L/\partial\theta\) for every parameter \(\theta\) in a network with \(L\) layers, by applying the chain rule from the loss backward to the inputs, reusing intermediate derivatives instead of recomputing them per-parameter (which would be exponential). It is the engine that makes gradient descent (below) tractable for millions/billions of parameters. (Ch. 18)
  • Batch normalization — a layer that re-centers and re-scales its input's activations to zero mean and unit variance per mini-batch (then applies a learned scale \(\gamma\) and shift \(\beta\)), which stabilizes and speeds up training by keeping activation distributions from drifting layer to layer ("internal covariate shift"). At inference time it uses a running average of mean/variance collected during training rather than the current (possibly batch-size-1) batch statistics — a common production bug is forgetting to switch the layer to eval mode, which silently reverts to batch statistics on a single request. (Ch. 18)
  • Batch size / epoch / iteration — a batch is the subset of training examples used to compute one gradient update; an iteration (or "step") is one such update; an epoch is one full pass over the entire training set. With 10,000 examples and batch size 100, one epoch is 100 iterations. Larger batches give a less noisy gradient estimate (lower variance) but each step is more expensive and, empirically, very large batches can generalize slightly worse without a matching learning-rate increase. (Ch. 18)
  • Bayesian optimization — a hyperparameter-search strategy that fits a probabilistic surrogate model (typically a Gaussian Process) over "hyperparameters → validation score" observed so far, then picks the next configuration to try by maximizing an acquisition function (e.g. Expected Improvement) that trades off exploring uncertain regions against exploiting known-good ones. It needs far fewer trials than grid or random search to reach a good optimum, which matters when each trial is a multi-hour training run. (Ch. 16, Ch. 17)
  • BeeGFS — an open-source parallel filesystem (originally Fraunhofer, now maintained by ThinkParQ) with an architecture similar in spirit to Lustre's (below) — separate metadata servers and storage servers, data striped across the latter, so the striping intuition transfers directly — but deliberately built for simpler installation and day-to-day operation than Lustre, at some cost in ecosystem maturity and enterprise feature depth. The typical home for a 50-300 GPU cluster standing up its first parallel filesystem without a dedicated Lustre specialist on staff. (Ch. 16b)
  • Bellman equation — the recursive identity that ties the value of a state to the value of its successor states, the foundation of dynamic programming and of almost every RL algorithm: $\(V^\pi(s) = \mathbb{E}_{a\sim\pi}\!\left[R(s,a) + \gamma\, V^\pi(s')\right].\)$ It says: the value of being in state \(s\) under policy \(\pi\) equals the immediate reward plus the (discounted) value of wherever you land next. Q-learning and value iteration are both fixed-point iterations on this equation. (Ch. 22)
  • Bias–variance decomposition — the identity that splits a model's expected test error into three additive terms: \(\text{Error} = \text{Bias}^2 + \text{Variance} + \text{Irreducible noise}\). High bias is underfitting (the model class is too simple to capture the pattern — e.g. a line through curved data); high variance is overfitting (the model chases noise specific to the training sample and won't generalize). Regularization (below) and more data both primarily attack variance; a bigger/more expressive model attacks bias. (Ch. 13)
  • Blue-green / Canary deployment — two zero/low-downtime release strategies. Blue-green keeps two full environments ("blue" = current, "green" = new) and switches all traffic atomically once green is validated — instant rollback is just switching back. Canary instead routes a small percentage of live traffic (e.g. 5%) to the new version, watches error/latency metrics, then ramps up gradually — it limits the blast radius of a bad release rather than avoiding downtime. (Ch. 8)
  • Boosting (Gradient Boosting, XGBoost, LightGBM) — an ensemble technique that builds trees sequentially, each new tree trained to predict the current ensemble's residual errors (technically, the negative gradient of the loss with respect to the current predictions), so the ensemble as a whole keeps reducing training loss step by step. It typically outperforms a Random Forest on tabular data at the cost of being more sensitive to overfitting and requiring more careful learning-rate/early-stopping tuning. Contrast with bagging (Random Forest), which builds trees independently and in parallel on bootstrap resamples and reduces variance by averaging, without the residual-fitting step. (Ch. 17)
  • BPE (Byte-Pair Encoding) — the standard subword tokenization algorithm for LLMs: start from individual bytes/characters, repeatedly merge the most frequent adjacent pair into a new token, until a target vocabulary size is reached. This lets a fixed vocabulary (e.g. 50k tokens) represent any string, including misspellings and unseen words, by falling back to smaller pieces. (Ch. 19)
  • Broadcasting — NumPy/PyTorch/TensorFlow's rule for applying an elementwise operation between arrays of different shapes by virtually stretching the smaller one along size-1 axes, without copying memory. Adding a shape-\((3,)\) vector to a shape-\((4,3)\) matrix broadcasts the vector across all 4 rows; the rule that makes this legal is that trailing dimensions must match or be 1. (Ch. 11)
  • Bulkhead (isolation pattern) — named for a ship's watertight compartments: cap the concurrent capacity (a connection pool, a thread pool, a semaphore) one dependency may consume so that a single slow-but-not-failing dependency cannot starve every other caller sharing the same process. Distinct from a circuit breaker (above), which stops calling a dependency that is failing outright — a bulkhead is the correct defense for a dependency that is still succeeding, just slowly. (Ch. 16g)
  • Burst buffer / local NVMe cache — a fast, typically NVMe-based storage tier interposed between compute nodes and a shared parallel filesystem, absorbing short, intense I/O bursts (a checkpoint write, a dataset stage-in) at local-device speed instead of hitting the slower or contended shared backend at peak rate. Applied to reads, a local NVMe cache holds a job's working-set data after a one-time stage-in, so every subsequent epoch reads from node-local NVMe — multiple GB/s, zero network hop, zero metadata-server round trip — instead of re-issuing shared-filesystem reads every epoch. A 1 TB working set on a node with 7.68 TB of usable NVMe fits with 6.68 TB of headroom; a 50-epoch job then converts 50 rounds of shared-filesystem contention into 1. (Ch. 16b)
  • Cancellation propagation — ensuring that when a client disconnects mid-request, the work being done on its behalf actually stops, rather than continuing to run (and, for a paid LLM provider call, continuing to bill) against an ASGI server (above) that has already detected the disconnect. In Python this means the asyncio.CancelledError raised at the current await point must be allowed to propagate out of the streaming loop and close the underlying upstream connection — a bare except Exception: (or a bare except asyncio.CancelledError: pass) around that loop swallows the signal silently, and generation keeps running against the provider with nothing that crashes or logs an error. (Ch. 19a)
  • CAP theorem — a distributed data store facing a network partition must choose between Consistency (every read sees the most recent write, or an error) and Availability (every request gets a non-error response, possibly stale); you cannot have both during a partition (\(P\) is not really a choice — networks do partition). Postgres/etcd (used for Kubernetes' own cluster state, above) are CP: they refuse to serve a read/write rather than risk an inconsistent answer when a quorum is unreachable. DynamoDB and Cassandra default to AP: they keep answering during a partition and reconcile divergent replicas afterward (eventual consistency). Neither choice is "better" in the abstract — a payments ledger wants CP, a shopping-cart "add to cart" button wants AP. (Ch. 12)
  • cgroups (control groups) — the Linux kernel feature that limits and accounts for a process group's CPU, memory, I/O, and network usage. Docker/Kubernetes resource limits/requests are a thin, friendly API over cgroups; an OOMKilled container (below) is the kernel enforcing a cgroup memory limit. (Ch. 1, Ch. 4)
  • Change detection (hash-based vs. validator-based) — the two ways a connector decides whether a document is worth re-parsing. Validator-based trusts a source's own metadata (an HTTP ETag, Last-Modified, or an API's updated_at) and treats "the validator changed" as "the content changed," cheaply, via a conditional GET that transfers no body when nothing changed. Hash-based distrusts the source's bookkeeping entirely: it fetches the bytes and compares a SHA-256 digest to the one stored from the previous crawl, trustworthy by construction but always paying for the full fetch. A source that bumps updated_at on a metadata-only touch makes validator-based detection over-trigger; one whose validator silently stops updating makes it under-trigger — hash-based comparison, run as a tripwire on every fetch actually performed, catches both failure directions. (Ch. 19b)
  • Chaos engineering — deliberately injecting failure into a running system (kill a random pod, add latency to a network call, exhaust a disk) in production or a production-like environment, on a controlled schedule, to verify that the redundancy/failover you designed on paper actually works under real conditions rather than only in the architecture diagram. Netflix's Chaos Monkey (randomly terminating instances) is the canonical example; the discipline generalizes to "game days" that rehearse an incident response before a real one forces the rehearsal to happen live. (Ch. 10)
  • Checkpoint (agent) — a persisted snapshot of everything needed to resume an agent trajectory from exactly where it left off: current state, accumulated context, and step index, written after every transition (not on a timer, which loses whatever transition happened since the last save). Deterministic replay (below) reads a checkpoint's recorded OBSERVE results, never a live tool, so a debugging session reproduces exactly what the original run saw. (Ch. 19e)
  • Chunk overlap / stride — overlap (\(o\), tokens) is the span of text repeated at the end of one retrieval chunk and the start of the next; stride (\(t = c - o\), tokens, where \(c\) is chunk size) is how far the start position advances between chunks — the same knob from opposite ends. Overlap exists to bound the probability \(p \approx \max(0, (s-o)/(c-o))\) that an answer span of length \(s\) straddles a chunk boundary: \(p = s/c\) at \(o=0\), and \(p=0\) once \(o \ge s\). It is not free — it inflates storage and embedding cost by \(c/(c-o)\) over the no-overlap baseline (+14.3% at \(o=64\), \(c=512\)), paid once at embedding time and again on every query. (Ch. 19b)
  • Chunked prefill — splitting a long prompt's prefill pass into fixed-size chunks and interleaving each chunk with pending decode steps from other in-flight requests, instead of running the whole prefill as one uninterruptible block. Without it, a single 8,000-token prompt dropped into a continuous batch (above) can stall every other request's next token for the entire duration of that one prefill pass — chunking bounds the worst-case delay any decode step can be made to wait behind someone else's prefill to roughly one chunk's compute time. (Ch. 16e)
  • CIDR (Classless Inter-Domain Routing) — the notation for an IP address block by prefix length, e.g. 10.0.0.0/16 means the first 16 bits are fixed (the network part) and the remaining 16 bits are host addresses, giving \(2^{16}=65{,}536\) addresses. A /24 gives \(2^{8}=256\) addresses (254 usable after the network and broadcast addresses) — the size you'll see on most VPC subnets. (Ch. 2, Ch. 5)
  • Circuit breaker — a client-side pattern that stops calling a downstream dependency once it has failed too often recently, failing fast (locally, instantly) instead of piling up slow timeouts against a dependency that is already down. It behaves like a state machine: closed (calls flow normally, failures counted) → trips to open (calls fail immediately without even attempting the network call) after the failure rate crosses a threshold in a rolling window → after a cooldown, half-open (let a single probe call through) → closed again if it succeeds, back to open if it doesn't. Combined with backoff/retry (above), a circuit breaker is what prevents one slow dependency from cascading into a full outage by exhausting every caller's thread pool waiting on it. (Ch. 2)
  • cloud-init — the standard first-boot provisioning mechanism most cloud images ship with: it reads user-data (a cloud-config YAML file or shell script) exactly once, at initial boot, to set a hostname, write a handful of files, or install a fixed package list, then never runs again on that instance. It has no re-entry point and no idempotence contract worth the name — a one-line configuration fix means terminating and re-booting the instance from scratch (Immutable infrastructure, above), not re-applying a script; Ansible's push model (above) is the tool for changing an already-running fleet's configuration without replacing it. (Ch. 16f)
  • CloudWatch — AWS's native metrics, logs, and alarms service: every managed AWS resource (EC2, RDS, Lambda, ALB…) emits metrics here automatically, and you can define alarms that page you or trigger autoscaling when a metric crosses a threshold. It's the AWS-native alternative/complement to Prometheus + Grafana. (Ch. 5, Ch. 10)
  • CNI (Container Network Interface) — the plugin interface Kubernetes uses to give every pod a routable IP address and wire up pod-to-pod networking (Calico, Cilium, Flannel are common implementations); the choice of CNI also determines whether NetworkPolicy (below) is enforced at all. (Ch. 6)
  • Coalesced memory access — a GPU warp's 32 threads reading contiguous, aligned addresses so the memory controller serves them in a single transaction rather than one per thread. HBM moves data in fixed-size (typically 128-byte) segments per transaction, so a warp reading one contiguous 128-byte segment costs 1 transaction, 100% useful bytes; the same warp reading a strided column (stride ≥ the segment size) can cost up to 32 transactions for the same useful bytes — a 32× HBM traffic amplification for identical arithmetic. (Ch. 16a)
  • Cold start (serverless) — the four-step latency tax (queue → image pull/sandbox init → runtime/model load → execution) a request pays when it arrives to find no warm instance waiting. Under Poisson arrivals with rate \(\lambda\) and a keep-alive window \(T\), \(P(\text{cold}) = e^{-\lambda T}\) — one request every 10 minutes against a 5-minute keep-alive gives \(\lambda T = 0.5\) and a 61% cold rate, which a fleet-wide p50 dashboard typically hides. (Ch. 16g)
  • Compaction (agent context) — any mechanism that reduces the token footprint of an agent's accumulated context without discarding what the trajectory still needs: summarize a span of turns, selectively evict disposable ones, keep a model-written scratchpad, or externalize a bulky observation to durable storage behind a pointer. Firing every \(k\) steps and retaining a fraction \(\rho\) of context each time bounds the peak context size at \(C_{\text{peak}} = ko/(1-\rho)\) — a fixed ceiling, independent of trajectory length, versus an uncompacted trajectory's unbounded quadratic growth. (Ch. 19e)
  • ConfigMap / Secret (Kubernetes) — key-value objects mounted into pods as environment variables or files; a ConfigMap holds non-sensitive configuration, a Secret holds sensitive values (base64-encoded, not encrypted, by default — encryption at rest must be configured separately). Neither is versioned: changing one does not automatically restart pods that already read it. (Ch. 6)
  • Configuration drift — the gap that opens between a fleet's declared configuration (what the last playbook or manifest run says should be true) and its real, current state, caused by any change applied outside that tool of record — a manual SSH fix at 3 a.m., a hand-edited config file. Drift is silent by construction: nothing fails, no alert fires, until either the drifted value causes an incident or a scheduled ansible-playbook --check --diff run reports it. Distinct from data/concept drift (above), which is a shift in an ML model's input distribution, not a fleet's software state. (Ch. 16f)
  • Confused deputy — a program holding more privilege than its requester, tricked into spending that privilege on the requester's behalf; the term predates LLMs by decades (the original 1988 example is a compiler service that writes its billing log to a caller-supplied filename). An LLM agent holding one ambient service credential for every conversation is the identical shape: it cannot tell, from inside its own reasoning, whether an instruction came from an authorized operator or from a paragraph an attacker planted in a document weeks earlier. (Ch. 19f)
  • Confusion matrix — the 2×2 (or \(k\times k\)) table of predicted vs. actual class counts that every binary classification metric is computed from: True Positives (TP), False Positives (FP), False Negatives (FN), True Negatives (TN). Worked example: TP=40, FP=10, FN=20, TN=930 (a 1000-example imbalanced set) gives precision \(=40/(40+10)=0.80\), recall \(=40/(40+20)=0.667\) — see F1 / precision / recall below for the combined score. (Ch. 13)
  • Container — a process (or process group) isolated by Linux namespaces (its own view of PIDs, network, filesystem mounts) and resource-limited by cgroups; unlike a VM, it shares the host kernel, which is why containers start in milliseconds instead of tens of seconds. (Ch. 4)
  • Continuous batching — an LLM-serving scheduler policy that evicts a finished request and admits a waiting one at the very next decode step, instead of holding every slot in a batch idle until the batch's longest member finishes (static batching). It keeps the GPU's batch dimension full at all times and is the single largest throughput lever in modern serving engines (vLLM, TGI) over a naive for prompt in prompts: model.generate(prompt) loop, which is static batching by accident. (Ch. 16e)
  • Convolution — a sliding, weight-shared linear filter applied over a spatial (or temporal) input: the same small kernel of weights (e.g. 3×3) is applied at every position, which gives the network translation equivariance (a shifted input produces a correspondingly shifted output) and drastically fewer parameters than a fully-connected layer over the same input. It is the core primitive of CNNs for images and 1D-convolutional models for sequences. (Ch. 18, Ch. 20)
  • Cost per task resolved — the dollar cost of running one agent trajectory attempt divided by the fraction of attempts that actually succeed, \(\text{cost per attempt}/\text{success rate}\) — never cost per call alone. Two agents with identical per-call cost can have wildly different costs per resolved task if their success rates differ; a design change that raises cost per attempt but raises success rate by more (a retry policy, e.g.) can be a net win even though a per-call dashboard reports it as a regression. (Ch. 19e)
  • CPCV (Combinatorial Purged Cross-Validation) — a cross-validation scheme for time series/financial data that (a) never trains on data that comes after a test fold chronologically without purging the overlap, and (b) additionally embargoes a buffer of samples immediately around each test fold, because overlapping labels (e.g. a triple-barrier label spanning several future bars) leak information across adjacent folds. Plain k-fold CV on time series without purging/embargo silently overstates performance — the single most common cause of backtests that "worked" and then failed live. (Ch. 23)
  • Crawl contract — the explicit SLA a document-ingestion connector commits to for one source: a rate limit (requests/second it will never exceed), a revisit interval (the ceiling on staleness a pull connector can promise), and a dedup key (the stable field — a source's own URI or native ID, never a title or path — used to recognize "the same logical document across crawls"). Two engineers each independently tuning a crawler's cadence without a written contract produce, in order: a rate-limit ban, an unnoticed freshness regression, and duplicate documents under two different keys after a rename. (Ch. 19b)
  • Cross-entropy — the standard classification loss, measuring how surprised a predicted distribution \(\hat p\) is by the true label \(y\): $\(L = -\sum_{k} y_k \log \hat p_k.\)$ For a one-hot true label (only class \(k^*\) has \(y_{k^*}=1\)), this collapses to \(L = -\log \hat p_{k^*}\) — the loss only depends on the probability mass the model put on the correct class. Worked example: a 3-class model outputs \(\hat p = [0.1, 0.2, 0.7]\) and the true class is the third one; the loss is \(L = -\log(0.7) = 0.357\) nats. If the model had instead been confidently wrong (\(\hat p_{k^*}=0.01\)), the loss would jump to \(-\log(0.01)=4.605\) nats — cross-entropy punishes confident mistakes much more than cautious ones. (Ch. 13, Appendix D)
  • Cross-validation — estimating a model's generalization performance by repeatedly splitting the data into train/validation folds (commonly \(k=5\) or \(10\), "k-fold CV") and averaging the validation score across folds, so every example serves as validation data exactly once. For i.i.d. tabular data this is the default; for time series it must be replaced by walk-forward or purged/embargoed schemes (see CPCV) because plain k-fold shuffles time order and leaks the future into the past. (Ch. 13, Ch. 17)
  • CRD / Operator — a Custom Resource Definition extends the Kubernetes API with a new object kind (e.g. ExternalSecret); an Operator is a controller that watches instances of that kind and drives real-world state toward what they declare (a control loop, exactly like the built-in Deployment controller, but for a domain-specific concept). (Ch. 6, ESO — Ch. 9)
  • CUDA / GPU — CUDA is NVIDIA's parallel-computing platform/API that lets frameworks (PyTorch, TensorFlow) offload dense linear algebra (matrix multiplies, convolutions) to the GPU's thousands of simple cores, which is what makes training networks with billions of parameters feasible in practice — the same operation on a CPU's tens of cores would take orders of magnitude longer. GPU memory (VRAM), not compute, is usually the binding constraint on how large a model/batch you can train. (Ch. 18)
  • CVD (Cumulative Volume Delta) — a market-microstructure signal that running-sums (buy volume − sell volume) over time; a rising CVD alongside a flat or falling price is a classic divergence signal used in order-flow-based trading strategies. (Referenced in the book's trading case studies.)
  • Cypher — the declarative graph query language MATCH (a)-[:REL]->(b) patterns are written in: a pattern reads left to right like an arrow drawn on a whiteboard. Originated with Neo4j, standardized as openCypher, and spoken (with dialect differences) by Memgraph, Amazon Neptune, and Apache AGE. A relationship pattern quantified as [:TYPE*min..max] matches a variable-length path in one clause — the graph-native answer to depth-unknown traversal that a fixed-depth SQL join can't express without hand-chaining one join per hop. (Ch. 19d)

D–H

  • DAG (Directed Acyclic Graph) — a graph with directed edges and no cycles, the natural data structure for anything with ordering dependencies and no circular waits: Airflow/dbt pipeline steps, CI/CD job graphs, and a neural network's computation graph are all DAGs. "Acyclic" is what guarantees a valid execution order (a topological sort) always exists. (Ch. 8, Ch. 12)
  • DaemonSet (Kubernetes) — ensures exactly one copy of a pod runs on every (or every matching) node in the cluster — the standard pattern for node-level agents like log shippers, CNI plugins, or Prometheus node-exporter, as opposed to a Deployment's arbitrary-N-replicas-anywhere model. (Ch. 6)
  • Data augmentation — synthesizing additional, label-preserving training examples by transforming existing ones (random crop/flip/rotation/color-jitter for images; back-translation or synonym swap for text; time-warping or noise injection for signals) so a model sees more variation than the raw dataset contains, without the cost of collecting/labeling new data. It is one of the cheapest and most reliable regularizers for deep learning, effectively fighting the variance side of the bias–variance trade-off (above) by widening the training distribution rather than by penalizing weights. (Ch. 18, Ch. 20)
  • Data lineage — the traceable record of where a dataset or feature came from and every transformation applied to it (source table → cleaning step → join → feature → model input), essential for debugging a bad prediction back to its root data cause and for audits/compliance. Tools like dbt and OpenLineage capture this automatically from pipeline definitions. (Ch. 12)
  • Data warehouse vs. data lake vs. lakehouse — a warehouse stores structured, schema-on-write data optimized for SQL analytics (Redshift, Snowflake, BigQuery); a lake stores raw files of any format cheaply on object storage, schema-on-read (S3 + Parquet); a lakehouse (Delta Lake, Iceberg, Hudi) adds ACID transactions, schema enforcement, and time travel on top of lake storage, aiming to get warehouse guarantees at lake cost. (Ch. 12)
  • DCGM (Data Center GPU Manager) — NVIDIA's fleet-grade GPU telemetry daemon, queried via dcgmi or scraped as Prometheus metrics by dcgm-exporter. It exposes per-pipe activity counters (DCGM_FI_PROF_SM_ACTIVE, PIPE_TENSOR_ACTIVE, DRAM_ACTIVE) that nvidia-smi's single GPU-Util number cannot distinguish, plus hardware-fault counters (Xid errors, ECC double-bit errors) that catch a GPU degrading silently before it fails outright. (Ch. 16a)
  • DDP (DistributedDataParallel) — PyTorch's data-parallel training wrapper: every rank holds a full model replica, computes gradients on its own data shard, and a ring all-reduce (below) averages gradients across ranks before the optimizer step, so every replica stays bit-identical. DDP overlaps this all-reduce with backpropagation by firing gradient buckets as soon as they're ready, hiding communication behind compute rather than paying for it afterward. Worked example: a 700M-parameter model's 1.4 GB bf16 gradient tensor, all-reduced over 8 GPUs at 25 GB/s, takes \(2(N-1)S/(N\beta) \approx 0.098\,\text{s}\) — a figure that barely grows as \(N\) scales further. (Ch. 16d)
  • Decision tree / Random forest — a decision tree recursively splits the feature space on the attribute/threshold that most reduces impurity (Gini or entropy for classification, variance for regression) at each node; a single tree overfits easily. A Random forest trains many trees on bootstrap-resampled data and random feature subsets per split, then averages their predictions (bagging) — the randomization decorrelates the trees' errors, so the ensemble's variance drops roughly by a factor of the number of trees (for uncorrelated errors) without increasing bias. (Ch. 17)
  • Deflated Sharpe Ratio (DSR) — the Sharpe ratio (below) corrected for the fact that testing many strategy variants and reporting only the best one inflates the apparent Sharpe purely by selection — exactly the multiple-comparisons problem in statistics. DSR asks: given that you tried \(N\) independent variants, what Sharpe would you expect the best one to show by pure luck, and is your observed Sharpe meaningfully above that? A widely used back-of-envelope approximation for the expected maximum Sharpe under the null (no real skill) over \(N\) independent trials is \(\mathbb E[\max \widehat{SR}] \approx \sigma_{SR}\sqrt{2\ln N}\). Illustrative example: with \(N=50\) backtested variants and a per-trial Sharpe standard error \(\sigma_{SR}\approx 1\), the expected noise-only maximum is \(\sqrt{2\ln 50} = \sqrt{2 \times 3.912} = \sqrt{7.824} \approx 2.80\) — so an observed Sharpe of 1.2 across 50 variants is not distinguishable from noise, even though 1.2 looks respectable in isolation. The exact Bailey & López de Prado formula (which also accounts for skew and kurtosis of returns) is derived in Appendix D. (Ch. 23)
  • Degradation ladder — an ordered sequence of fallback tiers a service steps down through as budget, quality, or security signals cross their thresholds, from full functionality to a hard stop, each step-down a deterministic, automatic transition rather than a decision that waits for a human to notice a dashboard. Stepping back up requires an explicit, human-confirmed reset — an automatic climb the instant a metric dips back under threshold makes the system flap between tiers. (Ch. 19f)
  • Dependency injection (layered service architecture) — structuring a service as router → service → repository → external-client layers, where each layer receives its collaborators (a database session, an idempotency store, a provider client) as constructor/function arguments rather than importing and instantiating them itself. FastAPI's Depends() wires this at the request level; the practical payoff is testability — a test hands the service layer a fake provider client behind the same interface the real one implements, exercising real business logic against a controlled double instead of mocking method calls on a concrete implementation. (Ch. 19a)
  • Deployment (Kubernetes) — the controller that manages a ReplicaSet of identical pods and drives rolling updates/rollbacks: change the pod template's image tag, and the Deployment controller creates new pods and terminates old ones gradually according to its maxSurge/maxUnavailable strategy, keeping the service available throughout. (Ch. 6)
  • Deterministic replay (agent) — reproducing an agent trajectory's exact sequence of decisions by feeding the model, at each step, the exact recorded OBSERVE result the checkpoint log captured the first time — never re-calling the live tool, which can return a genuinely different answer if the underlying state (an order's status, say) changed since. This is what makes time-travel debugging — loading a checkpoint from an earlier step and re-running forward — trustworthy: the replayed run sees what the original run actually saw, not a fresh, possibly-different observation. (Ch. 19e)
  • Diffusion model — a generative model trained to reverse a fixed process that gradually adds Gaussian noise to data over \(T\) steps until it becomes pure noise; generation runs that reversal from pure noise back to a sample, denoising a little at each of the \(T\) steps. It currently produces the highest-fidelity images of any generative family, at the cost of many sequential forward passes to sample. (Ch. 21)
  • Discount factor \(\gamma\) — the per-step multiplier (\(0 \le \gamma < 1\)) that shrinks the value of future rewards in the return \(G_t = \sum_{k=0}^{\infty}\gamma^k R_{t+k+1}\), both for numerical convergence (an infinite undiscounted sum can diverge) and to encode a preference for sooner rewards. With \(\gamma=0.99\), a reward 100 steps away is worth only \(0.99^{100}\approx 0.366\) of its face value today; with \(\gamma=0.9\) the same reward is worth \(0.9^{100}\approx 0.0000266\) — nearly discarded. Choosing \(\gamma\) is choosing the agent's effective planning horizon. (Ch. 22)
  • Dispatcher (PyTorch) — the layer between a Python-level op call (torch.add(x, y)) and the actual kernel that runs it: given a tensor's dtype, device, and active mode (autograd, autocast, vmap…), the dispatcher walks a key stack and picks exactly one registered kernel, then hands control to the next key down until a concrete CPU/CUDA kernel executes. Understanding the dispatcher explains why the same Python line can silently hit a different kernel under torch.autocast or inside a torch.compile'd region than it does eagerly. (Ch. 16c)
  • Docker Compose — a tool/file format (docker-compose.yml) that defines and runs a multi-container application (app + database + cache, say) as a single unit on one host, wiring up a shared network and named volumes between them — the right tool for local development and single-host demos, not for production orchestration (that's Kubernetes' job). (Ch. 4)
  • Dockerfile — the declarative recipe (FROM, COPY, RUN, CMD…) that builds a container image layer by layer; each instruction creates a new, cached, read-only layer, which is why instruction order matters for build speed (put rarely-changing steps like dependency installation before frequently-changing steps like copying source code). (Ch. 4)
  • Drift (data / concept) — data drift is a shift in the input feature distribution (e.g. average transaction amount doubles); concept drift is a shift in the true relationship between inputs and target (the same input now implies a different label/outcome). Both degrade a deployed model's accuracy silently — no error is thrown, predictions just get worse — which is why production ML needs Population Stability Index (PSI, below), KL divergence, or Kolmogorov–Smirnov (KS) tests monitoring the live input distribution against the training distribution. (Ch. 16)
  • DNS (Domain Name System) — the hierarchical, cached name → IP-address resolution system; each record has a Time-To-Live (TTL) controlling how long resolvers may cache it, which is why lowering a TTL before a planned cutover (so the change propagates fast) is standard practice. (Ch. 2)
  • Dropout — a regularization technique that randomly zeroes a fraction \(p\) of a layer's activations during each training step (and rescales the rest by \(1/(1-p)\) to keep the expected sum constant), which prevents units from co-adapting to each other's specific quirks — a crude but effective approximation of training an ensemble of exponentially many thinned networks and averaging them. Dropout is disabled at inference time. (Ch. 18)
  • DVC (Data Version Control) — git-for-data: DVC replaces a large file/directory in git with a small pointer file (a content hash), and stores the actual bytes in a configured remote (S3, GCS, a plain server). dvc pull/dvc push sync data the same way git pull/git push sync code, so a git commit can pin an exact dataset version alongside the exact code version that produced a model. (Ch. 14)
  • Dynamic inventory (Ansible) — a plugin that builds an Ansible inventory by querying a live source (the aws_ec2 plugin calls the EC2 API, azure_rm calls Azure Resource Manager) instead of reading a static .ini/YAML file of hardcoded hostnames, and groups hosts automatically from resource tags — keeping the inventory correct as instances are created and terminated by autoscaling or Terraform (Ch. 7), without anyone hand-editing a host list. cache: yes with a short TTL avoids re-querying the cloud API on every single playbook run. (Ch. 16f)
  • Early stopping — a regularization technique that monitors validation loss during training and stops (or restores the best checkpoint) once it stops improving for a configured number of epochs ("patience"), rather than training for a fixed, arbitrarily chosen number of epochs. It directly targets the bias–variance trade-off (above): training loss keeps falling as the model increasingly memorizes the training set, but validation loss eventually turns upward — early stopping cuts training at that turning point, before the memorization phase dominates. (Ch. 18)
  • EC2 (Elastic Compute Cloud) — AWS's raw virtual-machine service: you pick an instance type (CPU/RAM/ GPU ratio), an AMI, and a network placement (VPC/subnet/security groups), and get a billed-by-the-second VM. It's the substrate under many higher-level AWS services (EKS worker nodes are EC2 instances, for example). (Ch. 5)
  • ECR (Elastic Container Registry) — AWS's managed Docker image registry; access tokens returned by aws ecr get-login-password are short-lived (typically 12 hours) and must be refreshed before any docker push/pull, and any Kubernetes imagePullSecret built from that token expires on the same schedule. (Ch. 5, Ch. 6)
  • EKS (Elastic Kubernetes Service) — AWS's managed Kubernetes control plane: AWS runs and patches the API server/etcd, you run and scale the worker nodes (EC2 or Fargate) that actually host your pods. (Ch. 5, Ch. 6)
  • efConstruction / efSearch (HNSW) — the two candidate-list widths behind HNSW's greedy search (see ANN, above), differing only in when the search they size runs. efConstruction governs the search each new vector's insertion performs to pick its edges — it shapes the graph itself, so raising or lowering it for an already-built index means rebuilding the graph from scratch. efSearch governs the same search at query time against an already-fixed graph — it changes nothing structural, so it is hot-tunable per request with zero rebuild cost, and it is the parameter a team reaches for first when a recall or latency SLO shifts in production. Both trade more distance computations for a better chance of finding the true nearest neighbors; \(M\), the graph's max-degree parameter, sits with efConstruction as a build-time-only choice. (Ch. 19c)
  • ELBO (Evidence Lower BOund) — the objective a VAE (below) maximizes as a tractable stand-in for the true (intractable) data likelihood \(\log p(x)\): $\(\text{ELBO} = \mathbb{E}_{q(z\mid x)}\!\left[\log p(x\mid z)\right] - D_{\mathrm{KL}}\!\big(q(z\mid x)\,\|\,p(z)\big).\)$ The first term rewards accurate reconstruction; the second penalizes the learned encoder distribution \(q(z\mid x)\) for straying from a simple prior \(p(z)\) (usually a standard Gaussian), which keeps the latent space smooth and sample-able. Worked example of the KL term alone: if \(q(z\mid x) = \mathcal N(0.5, 1)\) and the prior is \(p(z)=\mathcal N(0,1)\), the closed-form KL between two univariate Gaussians is \(D_{\mathrm{KL}} = \ln\frac{\sigma_0}{\sigma_1} + \frac{\sigma_1^2 + (\mu_1-\mu_0)^2}{2\sigma_0^2} - \frac12\); with \(\sigma_0=\sigma_1=1\) and \(\mu_1=0.5,\mu_0=0\), this is \(0 + \frac{1+0.25}{2} - 0.5 = 0.125\) nats — a small, tolerable penalty for an encoder that only shifted the mean slightly. (Ch. 21)
  • ELB / ALB (Elastic / Application Load Balancer) — AWS's managed load balancers: ALB operates at layer 7 (HTTP/HTTPS), routing by host/path and terminating TLS, and is what a Kubernetes Ingress on EKS typically provisions under the hood via the AWS Load Balancer Controller. (Ch. 5, Ch. 6)
  • Embedding — a dense, low-dimensional vector representation of a discrete object (a word, a user, an item) learned so that geometric proximity (cosine similarity or Euclidean distance) reflects semantic similarity. The famous example: \(\text{embedding}(\text{king}) - \text{embedding}(\text{man}) + \text{embedding}(\text{woman}) \approx \text{embedding}(\text{queen})\) — arithmetic on meaning. (Ch. 19)
  • Entity resolution — deciding which of a set of extracted or ingested mentions refer to the same real-world entity and collapsing each co-referent group into one canonical record (also called record linkage / deduplication elsewhere in the literature). Comparing every mention pair is \(O(n_m^2)\) — at \(n_m=500{,}000\) mentions that is \(1.25\times10^{11}\) pairs, hours of compute before any real scoring runs — so production pipelines block first: partition on a cheap, high-recall key (a normalized name prefix plus jurisdiction, say) and score pairs only within a block, then take the transitive closure of above-threshold pairs (connected components of the match graph) to catch chains a single pairwise comparison misses. The failure that matters most is not a missed merge (cheap: a follow-up query catches it) but an over-merge — two distinct entities silently combined under one node, with no error thrown and no record of which property came from which original mention. (Ch. 19d)
  • Entra ID (formerly Azure AD) — the tenant-wide identity provider every Azure subscription attaches to: roughly AWS IAM and Cognito combined into one directory, since it issues both human-user sign-in and workload-identity tokens. Azure RBAC role assignments resolve against Entra principals — users, groups, service principals, managed identities — the same way IAM policies resolve against IAM principals. (Ch. 16g)
  • Ensemble methods — combining several models' predictions to get a result more accurate/robust than any single member: bagging averages independently trained models to reduce variance (Random Forest), boosting sequentially corrects residual errors to reduce bias (XGBoost), and stacking trains a meta-model on the base models' outputs. (Ch. 17)
  • Episodic memory (agent) — a durable record of past, completed agent trajectories: what happened, what tools were called, and what the outcome was, kept separately from the current trajectory's working memory (above the current context window) so a new trajectory can recall a relevant precedent without replaying it step by step. Typically lives in a relational store indexed by structured fields (customer, ticket type, outcome). (Ch. 19e)
  • Error budget — the allowed amount of unreliability implied by an SLO: if the SLO is 99.9% availability over 30 days, the error budget is \(1 - 0.999 = 0.001\), i.e. \(0.001 \times 30 \times 24 \times 60 = 43.2\) minutes of downtime you're allowed to "spend" — on risky deploys, experiments, maintenance — before you must freeze releases and focus purely on reliability. (Ch. 10)
  • Errors as data (agent tools) — a tool failure returned to an agent loop as a normal, schema-shaped observation — the same path a success would take — instead of raised as an exception that kills the loop. A retryable field lets the loop decide automatically whether to retry (a timeout) or route to a different action (an order ID that will never resolve) without the model having to guess why a call failed. (Ch. 19e)
  • ESO (External Secrets Operator) — a Kubernetes operator that syncs secrets from an external secret manager (AWS Secrets Manager, Vault…) into native Kubernetes Secret objects on a refresh interval, so the source of truth lives outside the cluster and secrets are never committed to git manifests. (Ch. 9)
  • etcd — the distributed, consistent key-value store that holds all Kubernetes cluster state (every object's desired spec and current status); it is the single most critical component to back up, because losing etcd means losing the cluster's memory of what should be running. (Ch. 6)
  • Event loop / coroutine / await point — the single-threaded scheduler at the heart of Python's asyncio (and ASGI, above): it holds a set of suspended coroutines and, at each await point inside one of them, checks whether that coroutine's I/O is ready and, if not, switches to running another coroutine instead of blocking the thread. A coroutine (an async def function) is a resumable computation, not a thread — it only ever runs on the loop's single thread, and it only ever yields control at an await. The corollary is the mechanism's one hard rule: any code inside an async def that does not await — a synchronous requests.get, time.sleep, a non-async database driver — blocks the entire loop for its full duration, capping the whole worker's throughput at \(1/\text{call duration}\) regardless of how many clients are waiting. (Ch. 19a)
  • Exploration–exploitation trade-off — the fundamental RL dilemma: exploiting the best action known so far maximizes immediate expected reward, but exploring untried actions might reveal something even better and is necessary to ever discover the true optimum. Concrete mechanisms: \(\varepsilon\)-greedy (act randomly with probability \(\varepsilon\), greedily otherwise), or entropy bonuses added to the loss (as in PPO, below) that reward the policy for staying stochastic. (Ch. 22)
  • F1 / precision / recall — from the confusion matrix (above): precision \(= \text{TP}/(\text{TP}+ \text{FP})\) answers "of everything I flagged positive, how much was actually positive?"; recall \(= \text{TP}/(\text{TP}+\text{FN})\) answers "of everything actually positive, how much did I catch?"; F1 is their harmonic mean, \(F_1 = 2 \cdot \frac{\text{precision}\cdot\text{recall}}{\text{precision}+\text{recall}}\), which (unlike the arithmetic mean) heavily penalizes a model that is great on one and terrible on the other. Worked example (continuing the confusion-matrix entry, TP=40, FP=10, FN=20): precision \(=0.80\), recall \(=0.667\), so \(F_1 = 2 \times 0.80 \times 0.667 / (0.80+0.667) = 1.067/1.467 \approx 0.727\). (Ch. 13)
  • Fairshare (Slurm) — the component of a pending job's scheduling priority reflecting how much of a shared cluster's capacity the submitting user or account has already consumed recently, relative to their allocated share: lower recent usage raises priority, higher recent usage lowers it, with usage weighted to matter less the further in the past it was spent. A team that has run jobs continuously for two weeks sees new submissions start behind a team that barely touched the cluster in that window, even if both jobs arrive at the same instant — a gap that narrows as the heavy user's recent usage ages out. sshare -a shows every account's current score; sprio -j <jobid> breaks one pending job's composite priority into its components (age, fairshare, QOS, size). (Ch. 16b)
  • Fargate — a serverless container runtime for ECS/EKS: you specify CPU/memory and a container image, AWS provisions and manages the underlying compute with no node to patch or scale yourself, billed per vCPU-second and GB-second actually used. (Ch. 5)
  • Feature engineering — transforming raw data into the inputs a model actually consumes: encoding categoricals (one-hot, target/mean encoding), scaling numerics, extracting date parts, building interaction terms, or computing domain signals (e.g. a rolling volatility for a trading model). Even in the deep-learning era, feature engineering remains decisive for tabular data, where raw deep nets routinely lose to a well-engineered gradient-boosted tree. (Ch. 13, Ch. 17)
  • Feature flag — a runtime on/off (or percentage-rollout) switch read from configuration rather than hardcoded, letting a team merge and deploy code continuously while controlling exposure separately from deployment: a half-finished feature can sit dark in production behind a flag, and a risky change can be enabled for 1% of traffic, then 10%, then 100%, watching metrics at each step — the same progressive- exposure idea as a canary deployment (above), but controlled by a config toggle instead of a separate release/rollback. (Ch. 8)
  • Feature store — a centralized system (e.g. Feast) that computes, stores, and serves features consistently for both offline training (batch, joined over history) and online inference (low-latency, point-in-time-correct lookup), eliminating training/serving skew caused by two independently maintained feature pipelines silently drifting apart. (Ch. 16)
  • FID (Fréchet Inception Distance) — a generative-image-quality metric that runs both real and generated images through a pretrained Inception network, fits a Gaussian to each set's feature activations, and measures the distance between the two Gaussians (mean and covariance); lower FID means the generated distribution is closer to the real one, both in average appearance and in diversity. (Ch. 21)
  • Filter collapse — the failure mode where a metadata-filtered ANN query (above) returns far fewer results than requested, often zero, because the filter was applied after a fixed-size ANN search already committed to a top-\(k\) chosen with no knowledge of the filter (a post-filter). If the filter is uncorrelated with proximity to the query, each of the \(k\) candidates independently survives with probability \(s\) (the filter's selectivity), so \(\mathbb{E}[\text{survivors}]=k\times s\) and \(P(\text{zero survivors})=(1-s)^k\): at \(s=1\%\), \(k=10\), that is ≈90% of queries returning nothing, with no error anywhere — the query is syntactically valid. The fix is a pre-filter: restrict the ANN traversal itself to already-matching candidates (a bitmap-restricted search or a per-filter-value index partition), never widening the post-filter's LIMIT as a permanent workaround. (Ch. 19c)
  • Fine-tuning vs. prompt engineering vs. RAG — three ways to specialize a pretrained model to a task without training from scratch: fine-tuning further updates the model's weights on task-specific data (full or parameter-efficient, e.g. LoRA below); prompt engineering only changes the input text at inference time, no weights touched; RAG (below) supplies the model with retrieved documents as context instead of, or in addition to, either. They compose — you can RAG a fine-tuned model with a carefully engineered prompt template. (Ch. 19)
  • for_each vs. count (Terraform) — both create multiple resource instances from one block, but count addresses them by numeric index (aws_instance.this[0]), so removing an item from the middle of the list shifts every later index and can force destroy/recreate of unrelated resources; for_each addresses them by a stable map/set key (aws_instance.this["prod"]), so removing one item only affects that item's address. Prefer for_each whenever the collection isn't a fixed, ordered count. (Ch. 7)
  • FSDP (Fully Sharded Data Parallel) — PyTorch's native implementation of ZeRO-3 (below): parameters, gradients, and optimizer state are sharded across ranks, and each layer's full parameters are reconstructed on demand via an all-gather immediately before it runs forward or backward, then freed. This trades roughly \(1.5\times\) plain DDP's (above) communication volume for memory that no longer scales with GPU count — the mechanism that lets a model too large for one GPU's memory train at all. limit_all_gathers bounds how many units' parameters can be resident (prefetched) at once, trading some overlap for a firm memory ceiling. (Ch. 16d)
  • FSx for Lustre — AWS's managed Lustre offering: it provisions and operates the MDS/OSS/OST topology (see Lustre, below) without the customer running Lustre administration directly, billed per provisioned capacity and throughput tier. It links natively to S3, lazily loading objects from a bucket on first read and writing results back — a straightforward high-throughput scratch layer in front of an S3-resident dataset for an AWS-native GPU fleet. (Ch. 16b)
  • GAN (Generative Adversarial Network) — two networks trained against each other: a generator \(G\) maps noise to fake samples, a discriminator \(D\) tries to tell real from fake, and \(G\) is trained to fool \(D\). The minimax objective is $\(\min_G \max_D\; \mathbb{E}_{x\sim p_{\text{data}}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))].\)$ At the (theoretical) equilibrium, \(D\) outputs \(0.5\) everywhere — it can no longer distinguish real from fake at all — which is also why GAN training is notoriously unstable (it's a saddle-point search, not a simple minimization). (Ch. 21)
  • GenAI semantic conventions (OpenTelemetry) — a standardized set of gen_ai.*-prefixed span attribute names OpenTelemetry defines for LLM requests specifically (gen_ai.system, gen_ai.request.model, gen_ai.usage.input_tokens / .output_tokens, gen_ai.response.finish_reasons), so a trace produced against any provider carries the same attribute names for the same facts, letting one dashboard or query work across providers without a custom attribute schema per integration. (Ch. 19f)
  • GIN (Generalized Inverted Index) — a Postgres index type that builds one entry per distinct element a column decomposes into (a full-text lexeme, an array member), each pointing to a posting list of every row containing it — unlike a B-tree, which orders scalar values. GIN is the only index type that answers "which rows contain this element" (a tsvector @@ tsquery match, an array-overlap && predicate) without a sequential scan; dropping it doesn't break the query, it just removes the difference between a sub-millisecond filtered lookup and a full table scan on every request. (Ch. 19b)
  • GitOps — the practice of making a git repository of declarative manifests the single source of truth for a system's desired state, with a controller (ArgoCD, Flux) continuously reconciling the live cluster to match that repo — so a deployment is a git commit/merge, and git revert is your rollback mechanism, fully auditable in history. (Ch. 8)
  • Goodput — the fraction of served requests that meet a stated latency SLO (both TTFT and TPOT thresholds, below), as opposed to raw throughput (tokens/s or requests/s regardless of how late any of them arrived). A replica can report full throughput while every request blows past its SLO five-fold — goodput is the number that catches that failure and throughput alone hides. (Ch. 16e)
  • GPFS (Spectrum Scale) — IBM's parallel filesystem, commercially licensed and also marketed as Spectrum Scale, that distributes both data and metadata management across cluster nodes via a distributed locking/token scheme rather than Lustre's (below) dedicated single-MDS-per-namespace split, and natively exposes the same data as POSIX, NFS, SMB, and S3 from one backend. Picked over Lustre when the same storage estate also has to serve non-HPC enterprise workloads and the multi-protocol, licensed support is worth it against running two filesystems. (Ch. 16b)
  • gres (Generic Resource, Slurm) — Slurm's mechanism for tracking consumable, non-CPU, non-memory resources per node, of which GPUs are the dominant example in an ML cluster. --gres=gpu:N requests \(N\) GPU devices per node; --gres=gpu:a100:N narrows the request to a specific GPU type on a generation-mixed partition. A job that requests CPU and memory but forgets --gres gets scheduled onto a GPU node with zero GPUs attached to its cgroup — nvidia-smi inside the job reports no devices found, with no queue error, because as far as Slurm's placement logic is concerned the job asked for CPU and memory only. (Ch. 16b)
  • GPUDirect RDMA — lets a NIC read/write GPU memory directly over the network, skipping the CPU-memory bounce buffer a naive transfer would otherwise pay for twice (GPU→host, host→NIC). It's what lets an inter-node ring all-reduce approach InfiniBand's (below) raw link bandwidth instead of being capped by PCIe-to-host-memory copy throughput. (Ch. 16d)
  • GPUDirect Storage — a direct DMA path between storage (local NVMe or a networked filesystem client) and GPU memory, via the cuFile API, skipping the CPU's staging ("bounce") buffer a normal read() uses — the storage-side counterpart to GPUDirect RDMA (above), applied to storage reads instead of network transfers. Material at fleet scale (freeing CPU cycles a data-loading pipeline needs for decode/ augmentation once bandwidth demand climbs into the tens of GB/s), a rounding error at small scale where a conventional read() path is nowhere near saturated. Requires driver, filesystem-client, and hardware support all the way down the stack, and falls back silently to the bounce-buffer path when any link is missing — verify with gdscheck, don't assume a flag being set means the fast path is active. (Ch. 16b)
  • GQA / MQA (Grouped-/Multi-Query Attention) — attention variants that shrink the KV cache (below) by caching fewer key/value head projections than there are query heads. Plain multi-head attention (MHA) caches one KV head per query head (\(n_{kv}=n_{\text{heads}}\)); MQA caches a single shared KV head (\(n_{kv}=1\)); GQA groups query heads and caches one KV head per group (\(1 < n_{kv} < n_{\text{heads}}\)), the middle ground production LLMs (Llama-3, Mistral) actually ship. Since KV-cache bytes/token scale linearly in \(n_{kv}\) (below), an 8-head-group GQA model with \(n_{kv}=8\) against a 32-query-head MHA baseline caches exactly \(32/8=4\times\) less per token — the same weights footprint, 4× the concurrent-request capacity on identical hardware. (Ch. 16e, Ch. 19)
  • Gradient boosting — see Boosting, above.
  • Gradient clipping — capping the magnitude of the gradient vector before the optimizer step, either by clipping each element to a fixed range (rare) or, far more commonly, by rescaling the whole gradient vector so its L2 norm never exceeds a threshold \(c\): if \(\lVert g\rVert_2 > c\), replace \(g\) with \(g \cdot c / \lVert g\rVert_2\). This prevents a single unlucky batch (or the compounding multiplicative effect of many layers, in a recurrent network or a very deep one) from producing an exploding gradient that destroys the model's weights in one step. Worked example: \(g=[3,4]\) has norm \(\sqrt{3^2+4^2}=5\); clipping to \(c=1\) rescales it to \([3,4]\times(1/5)=[0.6,0.8]\), same direction, unit norm. Standard practice for training Transformers and RNNs; rarely needed for well-behaved CNNs with batch normalization. (Ch. 18)
  • Gradient descent — the core optimization loop of nearly all of machine learning: repeatedly step the parameters \(\theta\) a small amount in the direction that most decreases the loss \(L\), $\(\theta \leftarrow \theta - \eta \nabla_\theta L,\)$ where \(\eta\) (the learning rate) controls the step size and \(\nabla_\theta L\) is the gradient (vector of partial derivatives) of the loss with respect to every parameter. Worked example: minimize the toy loss \(L(\theta) = (\theta - 3)^2\) starting at \(\theta_0 = 0\) with \(\eta = 0.1\). The gradient is \(\nabla L = 2(\theta - 3)\); at \(\theta_0=0\), \(\nabla L = -6\), so \(\theta_1 = 0 - 0.1\times(-6) = 0.6\). At \(\theta_1=0.6\), \(\nabla L = 2(0.6-3)=-4.8\), so \(\theta_2 = 0.6 - 0.1\times(-4.8) = 1.08\). Each step overshoots less and less, converging geometrically toward the true minimum \(\theta^* = 3\). In practice, plain gradient descent is replaced by an optimizer — see the entry below — that adapts the step per parameter. (Ch. 13, Appendix D)
  • Grafana — the open-source dashboarding front-end typically paired with Prometheus (query via PromQL, below) or any other metrics/logs/traces backend, used to build the RED/USE dashboards described in observability (below). (Ch. 10)
  • GraphRAG — a retrieval design that builds a hierarchy of LLM-written natural-language summaries over a knowledge graph's communities ahead of query time, then retrieves from those summaries instead of, or alongside, raw text chunks. Local search anchors to one or a few seed entities and reads their nearby leaf-community summaries — a bounded, cheap context; global search sweeps every summary at a coarse level of the hierarchy to answer a broad, corpus-spanning question no single entity anchors, and costs an order of magnitude more tokens per query as a direct result. The hierarchy's construction cost only earns its keep when a meaningful share of production queries are genuinely global; a corpus answering almost entirely entity-anchored questions doesn't need it built at all. (Ch. 19d)
  • Groundedness — the fraction of a generated answer's claims that are supported by the retrieved context actually provided to the model for that specific call, not by the model's general training knowledge and not by a claim that merely sounds plausible. A response can be fluent and confident while completely ungrounded if the model answered from parametric memory instead of the chunks it was actually given. (Ch. 19f)
  • Handler (Ansible) — a task that only runs when explicitly notify-ed by another task that reported changed, and even then only once per play no matter how many tasks notify it (deduplicated), and only at the play's end by default, after every task has run (flush_handlers, below, forces it earlier). The canonical use is "restart the service" after a config file changed — restarting is not itself idempotent (it disrupts a running process every time), so a handler defers and coalesces that action instead of firing it once per matching task. (Ch. 16f)
  • HBM (High Bandwidth Memory) / the memory wall — the stacked DRAM soldered onto a GPU package (1.55 TB/s peak on an A100 80GB SXM), the only path a kernel has to its input/output data. The memory wall is the structural, worsening gap between compute-throughput growth (~3× per hardware generation) and HBM-bandwidth growth (~1.6× per generation): the ratio between them widens roughly 12× over four generations, which is why an ever-larger share of workloads become memory-bound — capped by HBM bandwidth, not by how many FLOP/s the chip can issue — on each new GPU generation. (Ch. 16a)
  • Helm — the package manager for Kubernetes: a "chart" bundles a set of templated manifests plus a values.yaml of overridable parameters, so helm install myapp ./chart -f prod-values.yaml deploys a whole parameterized application (Deployment, Service, Ingress, ConfigMap…) in one command, and helm upgrade/helm rollback version the release as a unit. (Ch. 6)
  • HOT (Heap-Only Tuple) update — a Postgres UPDATE optimization that writes the new row version on the same heap page as the old one and skips writing any new index entries, applying only when no indexed column changed and the page has room. When it applies, an UPDATE costs one page rewrite; when it doesn't (any indexed column changed), every index on the table gets a new entry on top of the heap write. A row whose frequently-updated columns are all indexed — a chunk's text, content_hash, and embedding in a retrieval pipeline — disqualifies itself from HOT by construction. (Ch. 19b)
  • HPA (Horizontal Pod Autoscaler) — see Autoscaling, above; specifically the controller that scales a Deployment/StatefulSet's replica count up or down based on an observed metric (commonly CPU utilization, but any custom metric via the metrics API), keeping it near a target value you configure. (Ch. 6)
  • HuggingFace Hub — the git-backed model/dataset registry (huggingface_hub) that transformers pulls from by revision (a commit SHA or tag, never a mutable "latest"). Two pitfalls dominate in production: an unpinned revision means a repository maintainer's silent re-upload changes what your pipeline downloads on the next run with no error and no changelog, and trust_remote_code=True on an unfamiliar repository executes that repository's arbitrary Python at load time — audit the code before setting it, the same discipline as reviewing a dependency before adding it to a lockfile. (Ch. 16c)
  • Hyperparameter — any configuration value chosen before training that is not learned from data by gradient descent itself (learning rate, batch size, number of trees, regularization strength, network depth). Tuned via grid search (exhaustive), random search (usually more efficient per trial than grid for the same budget), or Bayesian optimization (above, most sample-efficient). (Ch. 13, Ch. 16)

I–O

  • IAM (Identity and Access Management) — AWS's system for defining who (users, roles, services) can do what (actions) on which resources (ARNs), expressed as JSON policy documents attached to principals. Nearly every AWS security incident traces back to an overly broad IAM policy ("Action": "*", "Resource": "*") rather than a broken encryption algorithm. (Ch. 5)
  • IaC (Infrastructure as Code) — describing infrastructure (networks, compute, IAM, DNS) in a declarative, version-controlled language (Terraform's HCL, CloudFormation, Pulumi) instead of clicking through a console, so infrastructure changes go through the same review/diff/history discipline as application code. Terraform's plan step is IaC's core safety net: it shows exactly what will change before anything is touched. (Ch. 7)
  • Idempotency — an operation that produces the same end state no matter how many times it's applied; PUT /users/42 {"name": "Ana"} is idempotent (repeating it changes nothing further), POST /users {"name": "Ana"} typically is not (each call creates a new user). Terraform apply and Kubernetes apply are both designed to be idempotent — safe to re-run after a partial failure or a retry. (Ch. 2, Ch. 7)
  • Idempotency key — a client-supplied token attached to a request so that a retried request (after a dropped connection, a client timeout) is served from a dedup store instead of triggering the underlying operation a second time. This matters more for a non-deterministic, billed AI generation than for a classic idempotent PUT (above): a naive retry against an LLM provider produces a second, different, separately billed generation for what the caller believes is one logical request, so the key must be checked and reserved in the repository layer before the provider is ever called, not after. (Ch. 19a)
  • Image / layer (Docker) — an image is an ordered stack of read-only filesystem layers plus metadata (entrypoint, exposed ports, env); layers are content-addressed and cached, so two images sharing a base layer (e.g. the same python:3.12-slim) don't duplicate that layer's storage or transfer time. (Ch. 4)
  • Immutable infrastructure — the practice of never patching a running server/container in place; instead, build a new image/AMI with the change baked in and replace the old instance wholesale. This eliminates configuration drift (two "identical" servers that silently diverged after years of manual patches) and makes rollback trivial (redeploy the previous image). (Ch. 4, Ch. 7)
  • Index-free adjacency — the storage design that makes a graph database a distinct category of engine rather than a relational query-planner trick: each relationship is a physical record holding direct pointers to its two endpoint nodes, and each node holds a pointer to its own first relationship record. Traversing to a neighbor costs one pointer dereference, \(O(1)\), regardless of table size — against a relational join's \(O(\log N)\) B-tree probe per edge. The two costs grow at the same geometric rate in hop count; what differs is the per-edge constant, which is why the graph's advantage is asymptotic in hop count and not guaranteed at a single hop, where a well-tuned index often still wins outright. (Ch. 19d)
  • Inference (batch vs. online) — batch inference scores a large dataset all at once on a schedule (nightly churn scores for every customer); online inference serves one prediction per incoming request with a latency budget (a fraud check on a live payment). The two impose very different infrastructure: batch optimizes for throughput (a big Spark/Ray job), online optimizes for p99 latency (a warm, autoscaled model server). (Ch. 16)
  • InfiniBand — a purpose-built, lossless, RDMA-native network fabric for HPC/AI clusters; a single HDR rail delivers on the order of 25 GB/s of effective, GPU-usable bandwidth once GPUDirect RDMA (above) is enabled, versus roughly 600 GB/s for intra-node NVLink (below) — the reason a ring all-reduce's bottleneck edge is usually the one crossing a node boundary, not the ones inside a node. See RoCE, the Ethernet-based alternative. (Ch. 16d)
  • Ingress (Kubernetes) — the object that routes external HTTP(S) traffic to internal Services by host and/or path, and typically also terminates TLS; it requires an Ingress controller (nginx-ingress, Traefik, AWS Load Balancer Controller) actually running in the cluster to do anything — the Ingress object alone is just a routing rule. (Ch. 6)
  • Init container / Sidecar container (Kubernetes) — two patterns for extra containers sharing a pod (above) with the main "app" container. An init container runs to completion before any app container starts (e.g. run a migration, wait for a dependency to be reachable) — if it fails, the pod doesn't start at all. A sidecar container runs alongside the app container for the pod's whole lifetime, sharing its network namespace and volumes (a log shipper tailing the app's log file, a service-mesh proxy intercepting all its traffic, a config-reloader watching a mounted ConfigMap) — the pattern that lets you bolt on cross-cutting infrastructure concerns without modifying the app container's image at all. (Ch. 6)
  • IoU (Intersection over Union) — the overlap metric for two bounding boxes (or segmentation masks): the area of their intersection divided by the area of their union, $\(\text{IoU} = \frac{|A \cap B|}{|A \cup B|}.\)$ Worked example: box \(A = [0,0,10,10]\) and box \(B = [5,5,15,15]\) (both axis-aligned, coordinates in pixels). Their intersection is the rectangle \([5,5,10,10]\), area \(= 5\times5=25\). Their union is \(|A|+|B|-|A\cap B| = 100+100-25=175\). So \(\text{IoU} = 25/175 \approx 0.143\) — well below the typical \(0.5\) threshold used to call a detection a "match," so this pair would be scored as a miss in most detection benchmarks. (Ch. 20)
  • Isolation level / Read Committed — an isolation level bounds which concurrency anomalies a transaction can observe. Postgres defaults to Read Committed: each statement inside a transaction gets its own fresh MVCC (below) snapshot, so two SELECTs in the same transaction can see different answers if another transaction committed in between (a non-repeatable read, allowed by design). Read Committed does guarantee that two concurrent UPDATEs against the same row never both silently succeed from a stale read — but only when the read-decide-write happens inside one statement Postgres itself can lock and re-check; a SELECT followed by a separate UPDATE gets none of that protection, which is the shape a lost update exploits. (Ch. 19b)
  • Jinja2 — the Python templating engine Ansible uses everywhere a value can vary: {{ var }} substitution inside a task or a config-file template, {% if %}/{% for %} control-flow blocks, and a library of filters (| default(...), | to_json, | product(...) for a nested loop's cross-product) applied with the pipe operator. A template: task renders a .j2 file through Jinja2 against the current host's facts and variables and writes the result to the target — the mechanism behind every fleet-specific config file a role generates. (Ch. 16f)
  • Job array (Slurm) — a single sbatch submission that creates \(N\) near-identical jobs, indexed 0..N-1, sharing one script but each seeing a distinct $SLURM_ARRAY_TASK_ID — Slurm's mechanism for submitting, tracking, and rate-limiting a hyperparameter sweep as one logical unit. --array=0-499%20 submits 500 tasks with a concurrency cap of 20 running at once: the cap trades wall-clock time against how much of a shared partition one sweep occupies, without changing the total GPU-hours consumed. (Ch. 16b)
  • Job / CronJob (Kubernetes) — a Job runs a pod to completion (retrying on failure up to a limit) for one-off/batch work, rather than keeping it running forever like a Deployment; a CronJob wraps a Job with a cron schedule (0 2 * * * = every day at 2am) to run it repeatedly. (Ch. 6)
  • JWT (JSON Web Token) — a compact, self-contained, digitally signed token made of three base64url-encoded, dot-separated parts (header.payload.signature): the header names the signing algorithm, the payload carries claims (sub = subject, exp = expiry, custom claims), and the signature (HMAC or RSA/ECDSA) lets any holder of the public key/shared secret verify the token wasn't tampered with without calling back to the issuer. This statelessness is exactly why OIDC (below) uses JWTs for identity tokens and why an API gateway can authorize a request by verifying a signature locally instead of a database round trip per request — at the cost that a JWT cannot be individually revoked before its exp without extra infrastructure (a deny-list), so short expiries plus refresh tokens are the standard mitigation. (Ch. 9)
  • Kafka — a distributed, partitioned, append-only log used as the backbone of streaming data pipelines: producers append events to a topic's partitions, consumers read at their own pace tracking an offset, and the log retains events for a configured window regardless of whether they've been consumed — which decouples producers from consumers entirely and lets multiple independent consumer groups replay the same stream. (Ch. 12)
  • KEDA (Kubernetes Event-Driven Autoscaling) — a Kubernetes autoscaler that scales on an external metric (a Prometheus query, a queue depth) instead of only CPU/memory, the mechanism a GPU-bound LLM serving pod needs: its process spends the overwhelming majority of its own CPU time waiting on the GPU, so CPU utilization stays low regardless of how saturated the replica actually is, and a plain HPA targeting CPU never fires. KEDA scaling on a serving engine's num_requests_waiting metric reacts to the real constraint instead. (Ch. 6, Ch. 16e)
  • Kernel — in machine learning, a kernel is a function \(k(x,x')\) that computes the inner product of two points as if they had been mapped into a much higher- (even infinite-) dimensional feature space, without ever computing that mapping explicitly (the "kernel trick"); this is what lets an SVM draw a nonlinear decision boundary in the original space while only ever solving a linear problem in feature space. In the operating system sense, the kernel is the privileged core of the OS that manages processes, memory, and hardware — cgroups and namespaces (containers' foundation) are kernel features. (Ch. 1, Ch. 17)
  • KL divergence — a (non-symmetric) measure of how different a probability distribution \(q\) is from a reference distribution \(p\): $\(D_{\mathrm{KL}}(p\,\|\,q) = \sum_k p_k \log\frac{p_k}{q_k}.\)$ It is zero iff \(p=q\) everywhere, and it is not a true distance (in general \(D_{\mathrm{KL}}(p\|q) \ne D_{\mathrm{KL}}(q\|p)\)). Worked example: \(p=[0.5,0.5]\), \(q=[0.9,0.1]\): \(D_{\mathrm{KL}}(p\|q) = 0.5\ln(0.5/0.9) + 0.5\ln(0.5/0.1) = 0.5\times(-0.588) + 0.5\times(1.609) = -0.294 + 0.805 = 0.511\) nats. This is exactly the quantity that appears in the VAE's ELBO (above) and in monitoring data drift (above) against a training-time reference distribution. (Ch. 16, Ch. 21, Appendix D)
  • Kubeconfig — the YAML file (~/.kube/config by default) holding cluster API endpoints, certificate authority data, and user credentials that kubectl reads to know which cluster to talk to and as whom; it can define multiple contexts (cluster+user+namespace triples) and switch between them with kubectl config use-context. (Ch. 6)
  • Kubernetes — the declarative container orchestration platform: you describe the desired state (how many replicas, which image, which resources) as objects, and a set of control loops continuously reconcile the actual cluster state toward it — self-healing (a crashed pod is recreated), rolling updates, service discovery, and autoscaling all fall out of this one reconciliation pattern applied to different object kinds. (Ch. 6)
  • Kueue / Volcano — two CNCF projects that bolt batch-scheduling semantics onto vanilla Kubernetes, which lacks them: jobs wait in an actual queue beyond immediate capacity rather than sitting Pending indefinitely, quota/fairsharing applies across teams, and — critically for multi-node training — gang admission only admits a Job's pods together once enough capacity exists for all of them, not piecemeal as individual GPUs free up. Without gang admission, a 4-pod/32-GPU job with only 3 nodes ready starts those 3 pods, which then block indefinitely inside torchrun's rendezvous waiting for the 4th rank — 24 GPUs allocated and idle for the full wait, \(24 \times 2\,\text{h} = 48\) GPU-hours wasted on zero training steps. Kueue is newer and more broadly adopted; Volcano (from Huawei) has a longer specific track record in ML gang-scheduling. Neither fully replaces Slurm's QOS/backfill maturity at scale. (Ch. 16b)
  • Kustomize — a template-free way to customize raw Kubernetes YAML: instead of parameterizing manifests with a templating language (Helm's approach), you write a base set of plain manifests plus small "overlay" patches per environment (dev/staging/prod), and Kustomize merges base + overlay at apply time. It ships built into kubectl (kubectl apply -k), which makes it a lighter-weight alternative to Helm when you don't need Helm's packaging/versioning/release-rollback machinery, only environment-specific tweaks (replica count, image tag, resource limits) on top of a shared base. (Ch. 6)
  • KV cache — per-request, per-layer storage of every past token's key and value projection vectors, held in GPU HBM so each LLM decode step can read the attended-to history instead of recomputing it from scratch every step. It grows by one token's worth of keys and values every decode step and is freed when the request completes. Its size per token is \(2 \times L \times n_{kv} \times d_{\text{head}} \times b\) bytes (\(L\) layers, \(n_{kv}\) cached KV heads — see GQA / MQA, above — \(d_{\text{head}}\) per-head dimension, \(b\) bytes/number); for a Llama-3-8B-shaped backbone (\(L=32\), \(n_{kv}=8\), \(d_{\text{head}}=128\), bf16) that is \(2\times32\times8\times128\times2 = 131{,}072\) B \(\approx\) 128 KiB/token, so a single 4,096-token-context request costs 512 MiB of GPU HBM — usually the resource a serving deployment runs out of before its compute. (Ch. 16e, Ch. 19)
  • Lakehouse — see Data warehouse vs. data lake vs. lakehouse, above.
  • Lambda (AWS) — a serverless compute service that runs your function code in response to an event (HTTP request, queue message, schedule) with no server to provision, billed per invocation and per GB-second of execution time; the trade-off versus a long-running server is per-invocation cold-start latency and a maximum execution duration. (Ch. 5)
  • Latent space — the (typically low-dimensional, continuous) space a generative model's internal representation lives in (a VAE's \(z\), a diffusion model's noise seed, a GAN's input noise); moving smoothly through latent space and decoding at each point produces smoothly morphing outputs, which is both a debugging tool (visualize what the model has learned) and a creative one (latent-space interpolation, arithmetic). (Ch. 21)
  • Learning rate schedule (warmup, cosine annealing) — the policy that varies the learning rate \(\eta\) (above, under Gradient descent) over the course of training rather than holding it fixed. Warmup linearly ramps \(\eta\) up from near zero over the first few hundred/thousand steps, avoiding a large, destabilizing update while the model's weights are still near their random initialization (critical for Transformers, above). Cosine annealing then decays \(\eta\) smoothly to (near) zero following \(\eta_t = \eta_{\min} + \tfrac12(\eta_{\max}-\eta_{\min})\big(1+\cos(\pi t/T)\big)\), where \(t\) is the current step and \(T\) the total number of steps — the cosine shape spends more steps near the peak learning rate (fast early progress) and tapers gently near the end (fine-grained convergence) compared to a linear decay. Worked example: \(\eta_{\max}=1\text{e-}3\), \(\eta_{\min}=0\), at the halfway point \(t/T=0.5\): \(\cos(\pi\times0.5)=\cos(90°)=0\), so \(\eta = 0 + 0.5\times(1\text{e-}3)\times(1+0) = 5\text{e-}4\) — exactly half the peak rate at the midpoint, as expected from the cosine shape. (Ch. 18, Ch. 19)
  • Least privilege — the security principle of granting a principal (user, role, service account) only the exact permissions it needs to do its job, nothing more — the single most effective, and most often skipped, mitigation against the blast radius of any credential leak or compromised workload. (Ch. 5, Ch. 9)
  • Lifespan (ASGI application) — the startup/shutdown hook an ASGI framework runs once per process, not once per request: code before the yield runs before the server accepts its first connection (the place to preload a model, open a connection pool), and code after the yield runs during shutdown (the place to drain in-flight requests). Nothing before the yield completes counts as the server being reachable at all — a slow preload (a 90 s model load, say) must be covered by a Kubernetes startupProbe (see "Probes," below) budgeted with real margin over that duration, or the default livenessProbe timing kills the still-loading pod in a crash loop. (Ch. 19a)
  • Little's law — see the worked example in "How to use this glossary," above: \(L=\lambda W\), the steady-state relationship between concurrency, arrival rate, and time-in-system, central to sizing worker pools and connection limits. (Ch. 2)
  • LLM-as-judge — using a second LLM call to score another model's output against a rubric (a short set of pass/fail or ordinal criteria), calibrated against a human-labeled sample before being trusted to score live traffic at scale. A judge calibrated once, at launch, drifts the same way any classifier's operating point drifts — a prompt change, a judge-model upgrade, or a shift in traffic composition can move its true agreement rate with nothing in a judge-only dashboard able to reveal it. (Ch. 19f)
  • LoRA (Low-Rank Adaptation) — a parameter-efficient fine-tuning technique that freezes the pretrained weight matrix \(W\) and learns only a low-rank update \(\Delta W = BA\) (with \(B \in \mathbb R^{d\times r}\), \(A \in \mathbb R^{r\times k}\), rank \(r \ll \min(d,k)\)) added at inference time; this cuts trainable parameters and optimizer memory by orders of magnitude versus full fine-tuning, at a small quality cost, and lets you swap task-specific LoRA adapters in and out of one frozen base model cheaply. (Ch. 19)
  • Loss function — the single scalar a training run is trying to minimize; it operationalizes "wrong" into a differentiable number so gradient descent has something to descend (mean squared error for regression, cross-entropy for classification, the GAN minimax objective, the RL policy-gradient objective — every learning algorithm in this book reduces to "define a loss, compute its gradient, step the parameters"). (Ch. 13)
  • mAP (mean Average Precision) — the standard object-detection benchmark score: for each class, plot precision against recall as the confidence threshold sweeps down, compute the area under that curve (Average Precision), then average AP across classes. "AP@0.5" means a predicted box only counts as a match if its IoU (above) with a ground-truth box exceeds 0.5. (Ch. 20)
  • Lustre — an open-source parallel filesystem, the dominant choice in traditional HPC and widely used for GPU training storage, built around separating where a file's data lives (the OST — Object Storage Target, physical storage on an OSS — Object Storage Server) from who tracks the namespace (the MDS — MetaData Server, which answers "where does this file live" once per open() and never serves a single byte of content itself). Aggregate read bandwidth is, to a first approximation, the sum of provisioned OSSs' individual bandwidth — the split that lets Lustre add throughput by adding storage servers instead of buying a faster single machine, unlike NFS. Striping (lfs setstripe --stripe-count N) spreads one file's bytes round-robin across \(N\) OSTs so reading it drives I/O against \(N\) OSSs in parallel; it is a pure win on few, large, sequentially-read files (sharded checkpoints, large WebDataset shards) and a pure RPC cost with zero throughput gain on small files that fit on one OST. Not retroactive: lfs setstripe only affects files created after it is set on a directory. (Ch. 16b)
  • Managed identity (Azure) — an identity Azure attaches directly to a resource (a VM, a Function App, an AKS pod via workload identity), so code on that resource acquires a token from the instance metadata endpoint with no secret ever stored — the direct Azure counterpart of an EC2 instance profile or EKS's IRSA. System-assigned identities share the resource's own lifecycle and die when it is deleted; user-assigned identities are standalone objects that survive a resource being recreated, the correct choice for anything (a scale set, a Key Vault access policy) whose lifecycle should outlive one resource. (Ch. 16g)
  • Management group (Azure) — the ARM hierarchy tier that sits above subscriptions and below the tenant root, the closest Azure analog to an AWS Organizations OU: Azure Policy assignments and RBAC role assignments made at a management group inherit down to every subscription (and every resource) nested beneath it. (Ch. 16g)
  • Metadata wall — the point at which a workload's real throughput ceiling is set not by a filesystem's aggregate data bandwidth but by the rate at which it can service namespace operations (open(), stat(), readdir()), because those are inherently bound to far fewer metadata servers than a bandwidth-provisioned filesystem has data servers. \(T_{open} = N_f \cdot t_{rtt} / W\) (\(N_f\) files, \(t_{rtt}\) metadata round-trip time, \(W\) parallel workers): \(10^6\) files at 5 ms/open with 16 workers costs 312.5 s spent entirely on open() calls before a single byte of training data is read — provisioning more bandwidth does nothing for this, because the bottleneck was never bytes per second. The structural fix is fewer, larger files (see WebDataset shard, below), the same argument Chapter 12b makes for a Parquet lake on an object store, applied to a POSIX-mounted training filesystem instead. (Ch. 16b)
  • MDP (Markov Decision Process) — the mathematical formalism underlying reinforcement learning: a tuple \((S, A, P, R, \gamma)\) of states, actions, transition probabilities \(P(s'\mid s,a)\), a reward function \(R(s,a)\), and a discount factor \(\gamma\) (above); "Markov" means the next state depends only on the current state and action, not on the full history — which is what lets the Bellman equation (above) be written recursively at all. (Ch. 22)
  • MFU (Model FLOPs Utilization) — achieved FLOP/s divided by a GPU's peak FLOP/s, the single number a design review asks for to separate "the job is slow because the model/schedule is inefficient" from "the job is slow because the hardware is saturated." A well-overlapped, well-batched dense-transformer training job typically reaches 40-55% MFU; a large gap below that points to the pipeline bubble (below), a comm-bound step, or an under-batched GEMM shape — not a hardware fault. (Ch. 16d)
  • MIG (Multi-Instance GPU) — NVIDIA's hardware-level GPU partitioning, available on A100/H100-class parts: it splits one physical GPU into up to 7 fully isolated instances, each with its own dedicated slice of SMs, tensor cores, L2 cache, and memory controllers. An A100 80GB's 1g.10gb profile gets 1/7 of compute and 10 of the 80 GB — a hardware, not just software, isolation guarantee, at the cost of a fixed per-instance ceiling and a full-GPU reset to reconfigure. (Ch. 16a)
  • MLflow — an open-source platform with four components used across this book: Tracking (log parameters/metrics/artifacts per run), Projects (packaged, reproducible runs), Models (a standard packaging format any serving tool can load), and the Model Registry (stage a specific run's model as Staging/Production, with lineage back to the exact run, code version, and data version that produced it). (Ch. 15)
  • Model parallelism (tensor / pipeline / sequence / expert / 3D) — the family of techniques that split a model itself across GPUs, used when the model — not the compute — no longer fits one card. Tensor parallelism splits an individual matmul's weight matrix, needs an all-reduce per layer per micro-batch, and therefore wants NVLink (below), not a cross-node fabric. Pipeline parallelism splits the model into sequential stages across GPUs, pipelining micro-batches through them; it pays a pipeline bubble (below) of idle time. Sequence parallelism splits activations along the sequence dimension to shrink per-GPU activation memory outside tensor-parallel regions. Expert parallelism (for Mixture-of-Experts models) places different experts on different GPUs and routes tokens between them with an all-to-all instead of an all-reduce. 3D parallelism composes tensor (innermost, NVLink-bound), pipeline (across nodes), and data parallelism (outermost) onto one physical cluster mesh. (Ch. 16d)
  • Molecule — the standard test framework for an Ansible role: a molecule/<scenario>/ directory pairs a driver (how disposable test instances are created — docker, spinning up throwaway containers, is the common $0 choice, no cloud spend) with a fixed stage sequence — dependency → lint → syntax → create → prepare → converge → idempotence → verify → destroy. The idempotence stage runs converge twice back to back and fails the build if the second pass reports any changed at all — the automated, merge-blocking version of the same three-run experiment Idempotency (above) describes doing by hand. (Ch. 16f)
  • MPI (Message Passing Interface) / rank / communicator / collective — the decades-old standard that gives HPC-native parallel processes a way to talk to each other once a scheduler has placed them. A rank is the unique integer ID (\(0\) to \(N-1\)) MPI assigns each process; a communicator (MPI_COMM_WORLD by default) is the named group of ranks a collective operation runs over — every rank calling it together, in the same order. Four collectives cover essentially all distributed-training traffic: broadcast (root sends the same data to everyone), reduce (every rank contributes, one combined result lands on the root), allreduce (like reduce, but the combined result lands on every rank — the data-parallel gradient-averaging collective), and alltoall (every rank sends a distinct chunk to every other rank, a full transpose — no summation, only redistribution; the pattern behind expert-parallel MoE token routing). Slurm implements a PMI/PMIx server inside srun, so an MPI job launched with srun --mpi=pmix gets rendezvous information directly from Slurm instead of a separate mpirun. (Ch. 16b, Ch. 16d)
  • Multi-stage build (Docker) — a Dockerfile with several FROM stages, where an early stage compiles/ builds (with all the heavy build tools) and a final stage COPY --from=<earlier-stage> only the compiled artifact into a minimal runtime base image — "build fat, ship lean," which shrinks the final image and its attack surface. (Ch. 4)
  • MVCC (Multi-Version Concurrency Control) — the mechanism behind Postgres's non-blocking reads: every row carries hidden xmin/xmax columns naming the transactions that created and superseded its current version, so an UPDATE never overwrites bytes in place — it inserts a new tuple and marks the old one dead, and a snapshot decides which version it sees by comparing its own visibility rules against xmin/xmax. This lets writers and readers run against the same table without blocking each other, at the cost that "delete the old version" is deferred to vacuum: every UPDATE leaves a dead tuple until autovacuum (above) reclaims it, and a long-open reader's snapshot (the xmin horizon) can pin dead tuples regardless of how busy that reader's query is. (Ch. 19b)
  • Namespace — in Kubernetes, a virtual sub-cluster used to scope names, apply resource quotas, and isolate RBAC (below) between teams/environments sharing one physical cluster. In Linux, a namespace is the kernel primitive that gives a process its own isolated view of some global resource (PIDs, network interfaces, mounts) — the actual mechanism containers (above) are built from. (Ch. 1, Ch. 6)
  • NAT Gateway — a managed AWS resource that lets instances in a private subnet initiate outbound connections to the internet (to pull packages, call an external API) while remaining unreachable from the internet inbound; billed per hour plus per GB processed, which makes it a common line item worth watching on an AWS bill. (Ch. 5)
  • NCCL (NVIDIA Collective Communication Library) — the library that implements GPU collectives (all-reduce, all-gather, reduce-scatter, all-to-all) as topology-aware rings or trees, not as a general network protocol; a communicator is the process group over which one NCCL call runs, and every rank in it maps 1:1 to a CUDA device. NCCL_DEBUG=INFO is the primary diagnostic: a via P2P/IB/GDRDMA line confirms the fast path, via NET/Socket reveals a silent fallback to plain TCP — a 10-20× slowdown with no crash. (Ch. 16d)
  • NetworkPolicy (Kubernetes) — a pod-level firewall: by default every pod can reach every other pod in the cluster, and a NetworkPolicy is an explicit allow-list (by pod/namespace label selectors and ports) that, once any policy selects a pod, switches that pod to default-deny for anything not explicitly allowed. Requires a CNI (above) that implements NetworkPolicy enforcement. (Ch. 6)
  • nlist / nprobe (IVF) — the two tuning knobs of an IVF (Inverted File) index (see ANN, above). nlist is the number of \(k\)-means clusters (Voronoi cells) the corpus is partitioned into at build time — a structural choice requiring a full re-clustering to change. nprobe is the number of those clusters actually scanned per query, \(1\le\text{nprobe}\le\text{nlist}\) — a query-time choice, hot-tunable with zero rebuild cost, trading recall (a true neighbor sitting in an unprobed cell is simply missed) against latency (more cells scanned, more vectors compared). IVF is usually paired with product quantization (below) for memory compression, the combination shipped as FAISS's IndexIVFPQ. (Ch. 19c)
  • Node affinity / anti-affinity / PodDisruptionBudget (Kubernetes) — three controls over where and how many pods may be evicted at once. Node affinity attracts a pod toward nodes matching a label (e.g. gpu=true), a softer, more expressive cousin of taints/tolerations (below) that expresses a preference from the pod's side rather than a repulsion from the node's side. Pod anti-affinity does the opposite between pods — e.g. "never schedule two replicas of this Deployment on the same node," so a single node failure can't take out every replica at once. A PodDisruptionBudget (PDB) caps how many pods of a set may be voluntarily evicted simultaneously (during a node drain for cluster upgrade, say) — e.g. minAvailable: 2 on a 3-replica Deployment guarantees the cluster autoscaler/upgrade process never drains the third replica until one of the first two is back, protecting availability during routine maintenance the way anti-affinity protects it during a hardware failure. (Ch. 6)
  • Normal equations — the closed-form solution that directly minimizes ordinary least-squares loss without any iterative gradient descent: $\(\theta = (X^\top X)^{-1} X^\top y.\)$ Worked example (simple linear regression, one feature plus intercept): three points \(x=[1,2,3]\), \(y=[1,3,3]\). Means: \(\bar x = 2\), \(\bar y = 2.333\). The slope is \(b_1 = \dfrac{\sum (x_i-\bar x)(y_i-\bar y)}{\sum (x_i - \bar x)^2}\); the deviations are \(x-\bar x = [-1,0,1]\) and \(y-\bar y = [-1.333, 0.667, 0.667]\), giving numerator \((-1)(-1.333)+0\times0.667+1\times0.667 = 1.333+0+0.667 = 2.0\) and denominator \(1+0+1=2.0\), so \(b_1 = 2.0/2.0 = 1.0\). The intercept is \(b_0 = \bar y - b_1\bar x = 2.333 - 1.0\times2 = 0.333\). The fitted line is \(\hat y = 0.333 + 1.0\,x\) — check: at \(x=1\), \(\hat y = 1.333\), close to the observed \(y=1\); the residual variance is what the model leaves unexplained. This closed form only exists for linear regression under squared loss; every other model in this book (logistic regression onward) needs iterative gradient descent instead. (Ch. 17)
  • NUMA (Non-Uniform Memory Access) — on a multi-socket or multi-GPU host, memory/GPU "distance" is not uniform: a CPU core or GPU reaching across a NUMA domain boundary to another domain's memory or device pays materially higher latency and lower bandwidth than reaching within its own domain. On a GPU node, co-scheduling a multi-GPU job's replicas across NUMA domains (rather than pinning them within one) can silently double a training step's synchronization time with no error anywhere in the stack. (Ch. 16a)
  • NVLink — NVIDIA's proprietary point-to-point GPU interconnect, roughly 19× the bandwidth of PCIe on the same host (≈600 GB/s aggregate per GPU vs. ≈32 GB/s per direction over PCIe). Multi-GPU training's all-reduce cost is dominated by whichever link the replicas' gradients actually cross, so placement relative to the NVLink fabric — not just GPU count — decides step time. (Ch. 16a)
  • Occupancy (GPU) — the fraction of a streaming multiprocessor's maximum resident warps that are actually resident, given a kernel's per-thread register usage and per-block shared-memory usage (whichever budget is more constraining sets the resident block count). High occupancy hides memory latency by giving the warp scheduler more warps to switch between while one is stalled — but it is not a throughput measurement: a kernel can sit at 100% occupancy and still be entirely bandwidth-starved. (Ch. 16a)
  • OIDC (OpenID Connect) — an identity layer on top of OAuth2 that adds a signed, verifiable identity token (a JWT) proving who authenticated, not just that an access token was issued; it's what Keycloak/oauth2-proxy use for human login (Ch. 9) and what GitHub Actions/GitLab CI use for keyless, short-lived AWS credentials via AssumeRoleWithWebIdentity (Ch. 8) — no long-lived AWS access key ever has to be stored in CI secrets. (Ch. 8, Ch. 9)
  • On-behalf-of (OBO) token — a request-scoped, short-lived, tenant- and action-bound token minted from an authenticated caller's session — the opposite on every axis of an ambient credential (a service-wide, long-lived credential the same for every request). Minted before the model reads a single token of context, it bounds what a hijacked model's proposed tool call can ever be authorized to do, regardless of what the model itself was steered into requesting. (Ch. 19f)
  • ONNX (Open Neural Network Exchange) — a framework-neutral graph format and an intermediate representation for trained models: exporting to .onnx freezes a model's computation graph so it can run under ONNX Runtime, TensorRT, or another engine without the original PyTorch/TensorFlow code present at inference time. Export is not lossless by default — dynamic control flow, unsupported ops, and shape assumptions can all silently change behavior — so production pipelines run a parity test (same input, compare original-framework and ONNX outputs numerically) as a release gate, not a one-time sanity check. (Ch. 16c)
  • OOMKilled — the Kubernetes/Docker status meaning the container's process was killed by the Linux kernel's OOM (Out-Of-Memory) mechanism for exceeding its cgroup memory limit — not a crash in the application's own code, and raising the memory limit (or fixing an actual leak) is the fix, not restarting the pod. (Ch. 1, Ch. 6)
  • Optimizer (SGD, Momentum, Adam, RMSProp) — the algorithm that turns the raw gradient into an actual parameter update, on top of plain gradient descent's \(\theta \leftarrow \theta - \eta\nabla L\). SGD with momentum accumulates a running average of past gradients (a "velocity") so the update keeps moving through small local bumps in the loss surface instead of stopping at every one. RMSProp and Adam additionally keep a per-parameter running average of squared gradients and divide the step by its square root, effectively giving each parameter its own adaptive learning rate — parameters with consistently large gradients get smaller effective steps, and vice versa. Adam (Momentum + RMSProp combined, with bias-correction terms) is the default choice for most deep-learning training in this book. (Ch. 18, Appendix D)
  • Orphaned vector — a retrieval chunk (and its embedding) that stays servable in the index after its source document has been deleted, because an incremental connector never received, and had no way to infer, a deletion signal for it: a document removed at the source simply stops appearing in "what changed since my cursor," the same structural blind spot query-based CDC has for deletes. The only fix is a periodic full-listing reconciliation that diffs the source's complete current listing against every stored document ID and tombstones (below) whatever's missing — the expected exposure window is half the reconciliation cadence, the worst case the full cadence. (Ch. 19b)
  • OWASP LLM Top 10 — OWASP's checklist of ten vulnerability categories specific to LLM applications (prompt injection, insecure output handling, training data poisoning, model denial of service, supply chain vulnerabilities, sensitive information disclosure, insecure plugin design, excessive agency, overreliance, model theft), useful for coverage — "did we forget a whole class of finding" — but not a substitute for likelihood-×-impact prioritization on a specific architecture. (Ch. 19f)

P–Z

  • PagedAttention — the KV-cache (above) allocator vLLM introduced: instead of reserving each request's cache as one contiguous block sized for its maximum possible context length (which wastes whatever fraction of that reservation the request never fills), it allocates cache in small fixed-size pages from a shared pool and tracks each request's pages in a per-request page table, the same idea an OS's virtual memory paging applies to process address spaces. This collapses internal fragmentation from routinely 20–40% of a naive contiguous allocator's reserved memory down to typically single-digit percent loss, turning that recovered memory directly into more concurrent-request capacity on the same GPU. (Ch. 16e)
  • Parquet — a columnar, typed, compressed binary file format: because values of the same column are stored contiguously, a query that only needs 3 of 50 columns reads a fraction of the bytes a row-oriented format (CSV, JSON) would require, and per-column encoding (dictionary, run-length) compresses far better than mixed-type rows. The default choice for any dataset that will be queried by column. (Ch. 11, Ch. 12)
  • Partition (Slurm) — a named, administrator-defined subset of a Slurm cluster's compute nodes, each carrying its own policy: which nodes belong to it, the maximum walltime a job in it may request, which QOS tiers (below) it accepts, and sometimes which users or accounts may submit to it at all. A typical ML platform carves partitions along hardware and policy lines — gpu-a100, gpu-h100, a debug partition with a short walltime cap so a sanity check never queues behind a week-long run. Picking the wrong partition either queues a job behind unrelated work or rejects it outright for exceeding that partition's walltime ceiling. (Ch. 16b)
  • PCA (Principal Component Analysis) — a dimensionality-reduction technique that finds the orthogonal directions (principal components) of maximum variance in the data, obtained as the eigenvectors of the data's covariance matrix, ordered by their eigenvalues (the variance each direction explains). Worked example: a 2-feature dataset with covariance matrix \(\Sigma = \begin{pmatrix}2 & 1\\1 & 2\end{pmatrix}\) has eigenvalues found from \(\det(\Sigma - \lambda I)=0 \Rightarrow (2-\lambda)^2 - 1 = 0 \Rightarrow \lambda = 3 \text{ or } 1\). Projecting onto just the first principal component (eigenvalue 3) retains \(3/(3+1) = 75\%\) of the total variance in a single dimension instead of two — the practical payoff of PCA: fewer dimensions for a controlled loss of information. It is computed in practice via the SVD (below) of the (centered) data matrix rather than by explicitly forming the covariance matrix, for numerical stability. (Ch. 17, Appendix D)
  • PEFT (Parameter-Efficient Fine-Tuning) / LoRA / QLoRA — the HuggingFace peft library's umbrella term for fine-tuning techniques that update a small fraction of a model's parameters instead of all of them. LoRA (above) is the most common PEFT method; QLoRA stacks one move on top: quantize the frozen base model's weights to 4-bit precision (nf4) while keeping the small trainable LoRA adapters in bf16. Worked example: a 7B-parameter base model's frozen fp16 weights cost \(7\times10^9 \times 2\,\text{B} = 14\) GB; at 4-bit nf4 that drops to \(7\times10^9 \times 0.5\,\text{B} = 3.5\) GB — small enough, plus a few hundred MB of LoRA adapter state, to fine-tune on a single 24 GB consumer GPU rather than an 80 GB data-center card. (Ch. 16c)
  • Perplexity — the standard language-model quality metric, defined as the exponential of the average cross-entropy (above) per token: \(\text{PPL} = \exp(H)\) where \(H\) is the mean cross-entropy in nats. Intuitively, perplexity is "the effective number of equally likely choices the model is confused among" at each position. Worked example: a model with average cross-entropy \(H = 1.5\) nats per token has perplexity \(\exp(1.5) \approx 4.48\) — as confused, on average, as if it were guessing uniformly among about 4–5 equally likely next tokens, even though the real vocabulary might have 50,000 entries. Lower is better; a perplexity near the vocabulary size means the model has learned almost nothing. (Ch. 19)
  • pgbouncer / transaction pooling mode — a lightweight process that multiplexes many client-side Postgres connections onto far fewer backend ones. In transaction pooling mode, a backend connection is checked out only for the duration of one BEGIN...COMMIT, released the instant it ends — collapsing "how many clients are connected" down to "how many transactions are genuinely concurrent right now" (sized by Little's Law, \(L = \lambda W\), over database transactions). It breaks anything relying on server-side state outlasting one transaction on the same physical backend, notably asyncpg's default prepared-statement cache, which must be disabled (statement_cache_size=0) behind it. (Ch. 19b)
  • Pipeline bubble — the fraction of a pipeline-parallel stage's wall-clock time spent idle while the pipeline fills (early micro-batches haven't reached later stages yet) and drains (late micro-batches have left earlier stages), equal to \((p-1)/(m+p-1)\) for \(p\) stages and \(m\) micro-batches per step. Worked example: \(p=4\), \(m=8\) gives \(3/11 \approx 27.3\%\) idle; doubling \(m\) to 16 (same \(p\)) drops it to \(3/19\approx15.8\%\) — the bubble shrinks toward zero as \(m\to\infty\), the opposite asymptotic behavior from the ring all-reduce's (below) convergence to a nonzero floor. (Ch. 16d)
  • Planner/executor (agent orchestration) — an orchestration topology where one planner step decomposes a task into a small set of sub-tasks up front, and one or more executor loops each run a sub-task's own state machine (above) independently. Cuts total trajectory token cost by shortening \(n\) per executor — the quadratic term in \(n\cdot b + o\cdot n(n+1)/2\) falls faster than the linear planning overhead added — but only pays off when the sub-tasks genuinely decompose ahead of time; a task whose branches depend on runtime results needs a graph topology instead. (Ch. 19e)
  • Playbook (Ansible) — a YAML file listing one or more plays, each mapping a host pattern (drawn from the inventory) to an ordered list of tasks, roles, and handlers. Running a playbook re-asserts every task's declared state against every matched host, in order, top to bottom; --check runs it in dry-run mode (report what would change, touch nothing) and --diff shows the exact content diff for any file a task would rewrite — the pairing this book's drift-detection cron relies on. (Ch. 16f)
  • Pod (Kubernetes) — the smallest deployable unit: one or more containers that always share the same network namespace (one IP) and can share volumes, scheduled together onto the same node and living/dying together. A Deployment doesn't manage containers directly — it manages Pods, which in turn hold containers. (Ch. 6)
  • Policy — in IAM, a JSON document listing allowed/denied actions on resources, attached to a user, role, or group. In reinforcement learning, a policy \(\pi(a\mid s)\) is the (possibly stochastic) mapping from states to a distribution over actions that the agent follows — the object that RL training is actually trying to improve. (Ch. 5, Ch. 22)
  • Positional encoding — since self-attention (above) treats its input as an unordered set of tokens (permuting the input permutes the output identically), a Transformer must inject explicit position information; the original design adds a fixed sinusoidal vector per position (varying frequency across dimensions) directly to each token's embedding before the first attention layer, so nearby positions get similar codes and the model can learn to attend by relative offset. (Ch. 19)
  • Positive predictive value (PPV) — \(P(\text{real positive} \mid \text{flagged})\), the same quantity as precision (see F1 / precision / recall, above) under the name used in detection and screening contexts. It is not a property of a detector alone: the same sensitivity and false-positive rate produce a very different PPV depending on how rare the real condition is (§D.4.3's base-rate fallacy). At realistic low prevalence, a detector's false-positive rate typically dominates PPV far more than its sensitivity does. (Ch. 19f)
  • PPO (Proximal Policy Optimization) — the most widely used policy-gradient RL algorithm in practice, which maximizes expected advantage (above) while clipping the probability ratio between the new and old policy to a band \([1-\varepsilon, 1+\varepsilon]\) (commonly \(\varepsilon=0.2\)) so a single update can't move the policy too far from the one that collected the data — trading a bit of per-step optimality for training stability. Worked example: if the raw ratio \(r(\theta) = \pi_{\text{new}}(a\mid s)/\pi_{\text{old}}(a\mid s) = 1.5\) (the new policy would have made this action 50% more likely) and \(\varepsilon=0.2\), PPO clips \(r(\theta)\) to \(\min(1.5, 1.2) = 1.2\) before multiplying by the advantage, capping how aggressively any single good (or bad) action can push the policy in one update. (Ch. 22)
  • Probes (readiness / liveness / startup) — Kubernetes health checks on a container: a startup probe gates when the other two probes even begin (for slow-booting apps); a liveness probe failing gets the container killed and restarted (it's deemed permanently stuck); a readiness probe failing only removes the pod from Service load-balancing (it's temporarily not ready, e.g. warming a cache) but does not restart it. Conflating liveness with readiness is a common cause of restart-crash-loops on pods that were merely slow, not broken. (Ch. 6)
  • Procedural memory (agent) — an agent's compiled how-to-act knowledge, distinct from raw facts (semantic memory, below) and from a log of past runs (episodic memory, above): a curated playbook ("when a ticket mentions a duplicate charge, check X before Y"), a routing rule baked into orchestration logic, or a fine-tuned behavior. Lives in a prompt template, a policy function, or model weights — never in the trajectory's own context. (Ch. 19e)
  • Product quantization (PQ) — a lossy vector-compression scheme used inside IVF-PQ and some HNSW variants (see ANN, above): split each \(d\)-dimensional vector into \(m\) contiguous subvectors, run \(k\)-means separately on each subvector position with \(2^b\) clusters, and store every vector as \(m\) codebook indices (\(b\) bits each) instead of \(d\times4\) raw fp32 bytes. Compression factor \(=\dfrac{d\times4}{m\times b/8}\): worked example, \(d=768\), \(m=96\), \(b=8\) gives \(768\times4=3{,}072\) raw bytes against \(96\times1=96\) coded bytes, a 32× reduction on the vector bytes alone. PQ is lossy — distances are computed against dequantized approximations, not the true vectors — which puts a hard ceiling on achievable recall that no amount of nprobe (above) can raise past. (Ch. 19c)
  • PromQL (Prometheus Query Language) — the query language for Prometheus's time-series metrics; e.g. rate(http_requests_total[5m]) computes the per-second average request rate over a trailing 5-minute window from a monotonically increasing counter — rate() on a counter, not a raw difference, is the idiomatic pattern because it correctly handles counter resets (process restarts). (Ch. 10)
  • Property graph — the graph data model nearly every production graph database (Neo4j, Memgraph, Apache AGE, Amazon Neptune's property-graph mode, ArangoDB) implements: nodes represent entities, relationships are directed, typed, binary connections between exactly two nodes, and both nodes and relationships may carry an arbitrary set of key–value properties plus one or more labels. Putting a property directly on a relationship (an ownership edge's pct, since) is the model's biggest ergonomic edge over an RDF triple (below), which needs an extra reification step (below) to attach anything to a bare (subject, predicate, object) triple. (Ch. 19d)
  • PSI (Population Stability Index) — a data-drift metric that buckets a feature's values (deciles, typically) and compares the proportion of examples in each bucket between a reference (training) period and a current (production) period: $\(\text{PSI} = \sum_i (p_i - q_i)\ln\frac{p_i}{q_i},\)$ the same functional form as a symmetrized KL divergence (above). A widely used rule of thumb: PSI \(< 0.1\) means no significant shift, \(0.1\)–\(0.25\) means moderate shift worth investigating, \(> 0.25\) means the feature distribution has changed enough to warrant retraining or a model audit. (Ch. 16)
  • PV / PVC / StorageClass (Kubernetes) — a PersistentVolume (PV) is a piece of real storage (an AWS EBS volume, say) provisioned in the cluster; a PersistentVolumeClaim (PVC) is a pod's request for storage matching some size/access-mode criteria, bound to a matching PV; a StorageClass parameterizes how PVs get dynamically provisioned on demand (which backend, which disk type) so you rarely create PVs by hand. (Ch. 6)
  • Pydantic v2 — the schema-validation library FastAPI uses to parse and validate a request body into a typed model, rewritten in Rust (pydantic-core) for a large speed-up over v1. Validation cost is real but small at the boundary and should not be repeated on data the service already trusts: re-validating each of 250 streamed output tokens through a full model, at roughly \(35\,\mu\text{s}\) per validation, adds \(250\times35\,\mu\text{s}\approx8.75\,\text{ms}\) of CPU work per request for data the service produced itself and never needed to distrust. (Ch. 19a)
  • Pyxis / enroot — a pair of NVIDIA-built tools that let an srun/sbatch job launch directly inside an OCI container image with GPU access, via a flag on the job submission itself (srun --container-image=nvcr.io#nvidia/pytorch:24.01-py3 ...), with no Kubernetes or any other orchestrator involved. Pyxis is a Slurm SPANK plugin, enroot a lightweight unprivileged container runtime; together they decouple "I want reproducible, portable container images" from "I need Kubernetes as my scheduler" — how many large-scale GPU cloud and neocloud providers actually run containerized training jobs under Slurm today. (Ch. 16b)
  • Q-learning — an off-policy, model-free RL algorithm that learns the optimal action-value function \(Q^*(s,a)\) directly via the Bellman optimality update \(Q(s,a) \leftarrow Q(s,a) + \alpha\big[r + \gamma\max_{a'}Q(s',a') - Q(s,a)\big]\), where \(\alpha\) is a learning rate; "off-policy" means it can learn the optimal policy's values while acting according to a different, more exploratory policy (e.g. \(\varepsilon\)-greedy) — a key reason it's compatible with replay buffers of old experience. (Ch. 22)
  • QOS (Quality of Service, Slurm) — a named policy object, independent of partition (above), that Slurm applies on top of whichever partition a job lands in — typically bounding max GPUs/jobs concurrently in use per user or per account, and optionally granting a priority boost relative to other QOS tiers. Partition answers "which nodes"; QOS answers "how much, and how urgently, within them," and the two combine multiplicatively: a gpu-a100 partition might accept both a normal QOS (32 GPUs/user, standard priority) and a high QOS (8 GPUs/user, priority boost) on the same physical nodes. (Ch. 16b)
  • Quantization — reducing the numerical precision used to store/compute a model's weights and/or activations (e.g. 32-bit floats down to 8-bit integers, or the 4-bit schemes common for serving large LLMs), trading a small, usually acceptable accuracy loss for large reductions in memory footprint and inference latency/cost. (Ch. 19)
  • RAG (Retrieval-Augmented Generation) — grounding an LLM's answer in retrieved documents rather than relying purely on what it memorized during pretraining: embed the knowledge base into vectors, embed the query the same way, retrieve the nearest documents (cosine similarity search), and stuff them into the prompt as context before generation — the standard mitigation for hallucination and for giving an LLM knowledge of private/recent data it was never trained on. (Ch. 19)
  • Rate limiting (token bucket) — protecting a service from being overwhelmed by capping how many requests a client (or the service as a whole) may make per unit time. The token bucket algorithm is the standard implementation: a bucket holds up to \(B\) tokens (the "burst" capacity), refills at a steady rate \(r\) tokens/second, and each incoming request consumes one token — if the bucket is empty, the request is rejected (HTTP 429) or queued. This allows short bursts up to \(B\) requests instantly while enforcing a long-run average rate of \(r\)/s, unlike a naive fixed-window counter which allows up to \(2B\) requests in a short window straddling two window boundaries. Worked example: \(B=100\), \(r=10\)/s; a client that has been idle accumulates up to 100 tokens and can fire 100 requests instantly, but is then limited to 10/s thereafter until the bucket refills — sized so a legitimate retry burst passes but a sustained scraping script does not. (Ch. 2)
  • RBAC (Role-Based Access Control) — Kubernetes' native authorization model: a Role (namespaced) or ClusterRole (cluster-wide) lists allowed verbs (get/list/watch/create/delete) on resource kinds, and a RoleBinding/ClusterRoleBinding grants that role to a user, group, or service account. Every pod that calls the Kubernetes API itself (an operator, a CI runner) does so as a service account bound to a role — least privilege (above) applies here exactly as it does in IAM. (Ch. 6)
  • RDF triple — a single fact in the Resource Description Framework, expressed as (subject, predicate, object), where subject and predicate are globally unique URIs; a graph is simply a set of triples, with no relationship-level place to hang a property. RDF Schema/OWL let a reasoner derive new triples and catch inconsistent data automatically, and every subject/predicate being a global URI is what lets independent datasets merge by simple set union — a property a property graph (above) does not have, and the reason regulatory/scientific ontologies (SNOMED CT, UniProt) are published as RDF. (Ch. 19d)
  • RDS (Relational Database Service) — AWS's managed relational database (Postgres, MySQL, etc.): AWS handles patching, automated backups/snapshots, and (with Multi-AZ) synchronous standby failover, in exchange for less low-level control than self-hosting the database on EC2. (Ch. 5)
  • Reciprocal rank fusion (RRF) — a way to merge two or more independently ranked retriever results (e.g. BM25 and dense ANN search, above) without ever comparing their raw scores, which are usually incommensurable (BM25 unbounded, cosine similarity bounded to \([-1,1]\)). For document \(d\), \(\text{RRF}(d)=\sum_i 1/(k_\text{rrf}+\text{rank}_i(d))\), summed over every ranked list \(d\) appears in; \(k_\text{rrf}=60\) is the near-universal default, a damping constant that keeps one retriever's single top pick from dominating the fused score. Because it uses only rank, not score, RRF rewards consensus: a document ranked reasonably well by every retriever beats one ranked first by only one of them and poorly by the rest. (Ch. 19c)
  • RED / USE methods — two complementary dashboarding philosophies: RED (for request-driven services) tracks Rate, Errors, Duration per endpoint; USE (for resources: CPU, disk, network) tracks Utilization, Saturation, Errors. A healthy on-call dashboard usually has one RED panel per service and one USE panel per critical resource, not an undifferentiated wall of graphs. (Ch. 10)
  • Region (AWS) — a fully isolated AWS geography (e.g. eu-west-3 / Paris) containing multiple Availability Zones (above); resources in one region are, by default, invisible to and unaffected by another region — replication across regions is always an explicit, deliberate choice. (Ch. 5)
  • Regularization (L1 / L2) — a penalty term added to the loss to discourage overly complex models and fight the variance side of the bias–variance trade-off (above). L2 (ridge) adds \(\lambda \sum_j \theta_j^2\), shrinking all weights smoothly toward zero without forcing any to exactly zero. L1 (lasso) adds \(\lambda\sum_j |\theta_j|\), which — because of the non-differentiable kink at zero — drives some weights to exactly zero, performing implicit feature selection. Larger \(\lambda\) means stronger regularization (more bias, less variance). (Ch. 13, Ch. 17)
  • Reification — promoting a relationship to a node in its own right so a fact that is really about the relationship — not about either endpoint alone — gets a stable identity other relationships can point at. An RDF triple (above) pays this cost by default for every relationship needing even one property: a bare triple carrying \(k\) properties reifies into \((3+k)\) triples. A property graph (above) pays it only when a relationship is genuinely n-ary (three or more real participants, as with a syndicated loan's several lenders) or needs attributes belonging to the relationship as a whole — a deliberate modeling choice there, not a structural tax. (Ch. 19d)
  • ReLU / sigmoid / softmax — the standard activation functions. ReLU \(f(x)=\max(0,x)\) is the default for hidden layers: cheap, and its constant gradient of 1 for \(x>0\) avoids the vanishing-gradient problem that plagued sigmoid-based deep nets. Sigmoid \(\sigma(x)=1/(1+e^{-x})\) squashes to \((0,1)\), used for a single-output binary-classification probability. Softmax \(\text{softmax}(x)_k = e^{x_k}/\sum_j e^{x_j}\) generalizes sigmoid to \(K>2\) classes, turning a vector of raw scores ("logits") into a valid probability distribution — this is exactly what feeds cross-entropy (above) at a classifier's output layer. (Ch. 18)
  • Rendezvous (torchrun / c10d) — the discovery protocol that lets a set of distributed-training worker processes find each other and elect a rank-0/master without any of them being told the others' addresses in advance: workers register at a shared endpoint (rdzv_endpoint) under a shared rdzv_id, and rendezvous completes once the expected worker count checks in. A stale rdzv_id left over from a botched restart is the most common cause of a job that never gets past init_process_group, with zero NCCL log lines anywhere. (Ch. 16d)
  • Residual connection — a shortcut \(y = f(x) + x\) that adds a layer (or block)'s input directly to its output, so the block only needs to learn the residual (the difference) rather than the full transformation from scratch; this keeps gradients flowing directly through the +x path during backpropagation, which is what made training networks with 50–150+ layers (ResNet) practical where plain stacks of layers had previously degraded with depth. (Ch. 18, Ch. 20)
  • ResNet / ViT — ResNet is a deep convolutional network built from residual blocks (above); ViT (Vision Transformer) instead slices an image into fixed-size patches, embeds each patch as a token, and applies a standard Transformer encoder (attention, above) over the patch sequence — trading the CNN's built-in translation-equivariance for the Transformer's greater capacity, at the cost of needing more training data to reach the same accuracy. (Ch. 20)
  • Retry budget — a cap on the cumulative fraction of a service's outbound traffic that may be retries (e.g. no more than 10% of calls to a provider may be retry attempts), independent of and complementary to backoff/jitter (above): backoff and jitter only spread retries out in time, they never reduce their total volume, so only an explicit budget prevents a struggling upstream from seeing amplified traffic during the very incident that is causing the failures. Worked example: at a per-attempt failure probability \(p=0.5\) and up to 2 retries (3 attempts total), the offered-load multiplier is \(1+p+p^2=1+0.5+0.25=1.75\times\) the client-visible arrival rate — precisely the moment a retry budget, not a longer backoff window, is what keeps that multiplier from compounding further. (Ch. 19a)
  • Reward model / RLHF — a reward model is a learned function (trained on human preference comparisons — "response A is better than response B") that scores a generated output the way a human rater would; RLHF (Reinforcement Learning from Human Feedback) then fine-tunes a language model with an RL algorithm (typically PPO, above) to maximize that learned reward model's score, which is how base LLMs are turned into helpful, instruction-following assistants. (Ch. 19, Ch. 22)
  • RFC 9457 (application/problem+json) — the IETF standard for a machine-readable HTTP error body: a JSON object with type (a URI identifying the problem category), title, status, and detail fields, so API clients can branch on a stable type instead of parsing a human-readable message string. The house rule this book applies on top of the RFC: never relay a provider's raw error body or status code straight through — translate it into your own service's problem+json contract, distinguishing retryable failures (rate limits, provider 5xx, timeouts) from terminal ones (the provider rejecting a request your own code built incorrectly, which no number of retries will fix). (Ch. 19a)
  • Ring all-reduce — the algorithm every production collective-communication library (NCCL, above) uses to average gradients (or any per-GPU tensor) across \(N\) GPUs arranged in a logical ring: a reduce-scatter phase (each GPU ends up holding one fully-reduced \(1/N\) slice) followed by an all-gather phase (every GPU collects every slice), \(2(N-1)\) steps total, each moving only \(S/N\) bytes. Total time is \(T = 2(N-1)S/(N\beta) \to 2S/\beta\) as \(N\to\infty\) — communication cost that stops growing with cluster size, the property that makes data-parallel training viable at scale at all, unlike a naive central-parameter-server design whose cost grows as \(O(N)\). (Ch. 16d)
  • RoCE (RDMA over Converged Ethernet) — runs the same RDMA semantics as InfiniBand (above) over standard Ethernet switches, at comparable achieved bandwidth but with a higher CPU/latency tax and a hard dependency on lossless-Ethernet configuration (PFC/ECN) that a misconfigured switch silently breaks, degrading transport to plain TCP with no error — only a 10-20× slowdown visible in NCCL_DEBUG=INFO as via NET/Socket. (Ch. 16d)
  • Rolling update — the default Kubernetes Deployment rollout strategy (above): replace old-version pods with new-version pods a few at a time (governed by maxSurge/maxUnavailable), keeping the service continuously available, as opposed to a "recreate" strategy that tears everything down before bringing the new version up. (Ch. 6, Ch. 8)
  • Rolling upgrade (Ansible: serial / max_fail_percentage) — the fleet-configuration counterpart to Kubernetes' Rolling update (above), but at the level of hosts a playbook targets rather than pods a Deployment manages: serial: N restricts a play to fixed-size batches of N hosts, running every task (and flushing handlers) to completion for one batch before the next batch starts, so at most N hosts are ever unavailable at once. max_fail_percentage sets the fraction of one batch allowed to fail before the whole rollout aborts, evaluated fresh per batch, never cumulatively. Sizing serial against an SLO's capacity floor \(N_{\text{fleet}} \geq C + k\) (\(C\) = minimum online nodes required, \(k\) = safety margin) sets the largest batch a fleet can absorb without breaching its own SLO mid-upgrade. (Ch. 16f)
  • Roofline model — the decision tool that classifies any kernel as compute-bound or memory-bound from two hardware limits and one property of the kernel: achievable throughput \(= \min(P_{\text{peak}},\ I \times B)\), where \(P_{\text{peak}}\) is peak compute (FLOP/s), \(B\) is peak memory bandwidth (byte/s), and \(I\) is the kernel's arithmetic intensity (FLOP/byte). The ridge point \(I_{\text{ridge}} = P_{\text{peak}}/B\) splits the two regimes: below it, more bandwidth (or more reuse, raising \(I\)) is the only lever; above it, only faster or lower-precision compute helps — more of the other resource does nothing. On an A100 80GB SXM, \(I_{\text{ridge}} \approx 201\) FLOP/byte. (Ch. 16a)
  • Route 53 — AWS's managed DNS (above) service, also supporting health-check-based failover routing and weighted/latency-based routing policies across regions/endpoints. (Ch. 5)
  • Row-level security (RLS) — a Postgres feature (CREATE POLICY after ALTER TABLE ... ENABLE ROW LEVEL SECURITY) that ties every row of a table to a session-level setting (current_setting('app.tenant_id'), say), making a tenant filter mandatory and enforced by the database itself on every query, regardless of whether the application code that issued the query remembered to add the WHERE clause. (Ch. 19f)
  • S3 (Simple Storage Service) — AWS's object storage: durable, effectively infinitely scalable key-value storage for arbitrary blobs, billed by GB stored, requests, and egress bandwidth; it is the substrate under most data lakes (above), Terraform remote state, and DVC/MLflow artifact storage in this book. (Ch. 5, Ch. 7, Ch. 12)
  • safetensors — the HuggingFace-designed serialization format for model weights that replaced pickle (.bin) as the transformers/peft default: a flat header describing each tensor's name, dtype, and shape, followed by raw tensor bytes, with no embedded executable code. Loading a .bin checkpoint means unpickling it, which can execute arbitrary code the file chooses to embed; loading a .safetensors file cannot, because the format has no opcode for "run this." It also loads faster in practice via mmap — the OS pages tensor bytes in on demand instead of deserializing the whole file up front. (Ch. 16c)
  • SAST / DAST — Static Application Security Testing scans source code (or dependency manifests) for known-vulnerable patterns/libraries without running the application; Dynamic Application Security Testing probes a running instance of the application (fuzzing inputs, checking headers) for exploitable behavior — the two catch different classes of issues and are typically both wired into CI/CD's security gate. (Ch. 8, Ch. 9)
  • Scale-to-zero — a serverless platform's ability to reduce a route's running instance count to zero when no traffic has arrived within its keep-alive window, billing nothing while idle in exchange for paying the next request a cold-start tax (above). The property that makes serverless cheap at low, bursty traffic and expensive in latency at exactly the same traffic shape. (Ch. 16g)
  • Seccomp — a Linux kernel facility that restricts a process to an explicit allowlist of system calls, refusing any syscall not on the list before user-space code (however malicious) ever gets a chance to run it. Paired with a read-only root filesystem and container isolation, it bounds what a compromised process can do to its host regardless of what code ends up executing inside it. (Ch. 19f)
  • Security Group vs. NACL — an AWS Security Group is a stateful firewall attached to an instance/ENI (an allowed inbound connection's return traffic is automatically allowed, no matching outbound rule needed); a Network ACL is a stateless firewall attached to a subnet (return traffic must be explicitly allowed by a separate rule) evaluated in numbered-rule order. Most designs rely on Security Groups for day-to-day rules and NACLs only for coarse subnet-wide deny rules. (Ch. 5)
  • Semantic memory (agent) — facts and documents about the world an agent operates in: policy text, product catalogs, a knowledge base — as opposed to episodic memory's record of what the agent itself did. Typically lives in the vector index a RAG pipeline already retrieves from, so an agent's semantic recall reuses the same retrieval infrastructure rather than a bespoke store. (Ch. 19e)
  • Service (Kubernetes) — a stable virtual IP address and DNS name (myapp.namespace.svc.cluster.local) that load-balances traffic across the current, ever-changing set of pods matching a label selector — the layer that lets other components address "the app" without ever needing to know an individual pod's (ephemeral) IP. (Ch. 6)
  • Sharpe ratio (and Sortino, Calmar) — the classic risk-adjusted return metric, $\(\text{Sharpe} = \frac{\mathbb E[r] - r_f}{\sigma_r},\)$ the mean excess return (over a risk-free rate \(r_f\), often approximated as 0 for short horizons) divided by the return's standard deviation. Worked example: a strategy has mean daily return \(\mu = 0.001\) (0.1%/day) and daily standard deviation \(\sigma = 0.015\) (1.5%/day); the daily Sharpe is \(0.001/0.015 = 0.0667\). Annualizing (assuming i.i.d. daily returns and 252 trading days) scales by \(\sqrt{252}\approx 15.87\): \(\text{Sharpe}_{\text{annual}} \approx 0.0667 \times 15.87 \approx 1.06\). Sortino replaces \(\sigma_r\) with downside deviation only (volatility from losing days doesn't hurt a strategy the way volatility from winning days does — Sortino doesn't penalize it). Calmar instead divides annualized return by maximum drawdown, directly answering "how much return per unit of the worst peak-to-trough loss experienced." A high Sharpe reported after testing many variants should always be checked against the Deflated Sharpe Ratio (above). (Ch. 23)
  • SLI / SLO (/ SLA) — a Service Level Indicator is the raw measured metric (e.g. "fraction of requests under 200ms"); a Service Level Objective is the internal target for that indicator (e.g. "99.5% of requests under 200ms, measured over 28 days"); a Service Level Agreement is the external, often contractual, commitment (typically looser than the internal SLO, leaving margin) with a business consequence (credits, penalties) if missed. The error budget (above) is derived directly from the SLO. (Ch. 10)
  • slurmctld / slurmd / slurmdbd — Slurm's three daemons, each owning a distinct concern. slurmctld is the controller daemon — one process (with an optional hot-standby backup) holding the cluster's entire live state and making every placement decision; its downtime freezes new scheduling activity cluster-wide, though already-running jobs are unaffected. slurmd is the node daemon — one process per compute node that reports resources, receives launch instructions, and actually fork()/exec()s and cgroups a job; its downtime loses only the jobs on that one node. slurmdbd is the database daemon — a separate, MySQL/MariaDB-backed process that persists job accounting history and stores QOS/association definitions, the data source fairshare (above) reads to compute priority; its downtime does not affect live scheduling, only historical sacct queries and fresh accounting. The operational summary: "losing a slurmd is a Tuesday; losing slurmctld is a page." (Ch. 16b)
  • SNS / SQS — Simple Notification Service is AWS's pub/sub fan-out (one message delivered to many subscribers — email, Lambda, SQS queues); Simple Queue Service is AWS's point-to-point queue (one message consumed by one worker among a pool), used to decouple producers from consumers and absorb load spikes. (Ch. 5)
  • Speculative decoding — trading spare compute for latency during LLM decode: a small, cheap draft model proposes \(\gamma\) tokens ahead, then the full target model verifies all \(\gamma\) in a single batched forward pass and accepts the longest correct prefix, discarding the rest. Because decode is memory-bandwidth-bound — the target model's own weight read dominates one accepted token's cost regardless of how many draft tokens were checked alongside it — verification is nearly free bandwidth-wise up to the point the acceptance rate \(\alpha\) drops low enough that most \(\gamma\)-token batches are wasted work — a workload-dependent break-even, not a fixed win. (Ch. 16e)
  • SPMD sharding (pjit / jax.jit with PartitionSpec) — JAX's model-parallel counterpart to pmap: instead of replicating the full weight matrix on every device and only splitting the batch (pmap's Single-Program-Multiple-Data-on-a-replica model), sharded jit partitions the weight matrix itself into slices distributed one per device via a PartitionSpec, and XLA inserts the collective communication (all-gather, all-reduce) each op boundary needs automatically. For \(N\) devices, pmap costs \(N\times\) the weight memory of one full replica; sharded jit costs \(1\times\) — a \(4\times\) weight-memory ratio for 4 devices, the difference between a 7B model's 112 GB fitting or not fitting when replicated across GPUs with less memory than that per card. (Ch. 16c)
  • SSE (Server-Sent Events) — a text-based, unidirectional streaming protocol layered directly on HTTP: the server keeps a response open and pushes data: ... lines as they become available, and the browser's native EventSource (or any streaming HTTP client) reads them incrementally, without the bidirectional complexity of WebSockets. It is the standard transport for token-by-token LLM output; a reverse proxy or ingress with output buffering enabled (the default in nginx and many ALBs) can silently collect the whole stream before forwarding it, turning a token-by-token response back into one all-at-once response with no error raised — disable buffering on the route explicitly and verify with curl --no-buffer against the proxy itself. (Ch. 19a)
  • Stationarity (and the ADF test) — a time series is (weakly) stationary if its mean, variance, and autocorrelation structure don't change over time; most classical forecasting models (ARIMA, above) and many statistical tests assume it. The Augmented Dickey-Fuller (ADF) test checks the null hypothesis "the series has a unit root" (is non-stationary); rejecting the null (typically at \(p<0.05\)) supports treating the series as stationary, while failing to reject usually means you should difference it (subtract each value from the previous one, the "I" in ARIMA) before modeling. (Ch. 23)
  • State (Terraform) — the JSON file (terraform.tfstate) mapping every resource block in your .tf code to the real-world resource ID it created, which is how terraform plan computes a diff between code and reality. It must live in a locked, remote backend (S3 + DynamoDB lock table, or Terraform Cloud) shared by the whole team — a local, unlocked state file is both a single point of failure and a race condition waiting to corrupt itself under concurrent applys. (Ch. 7)
  • StatefulSet (Kubernetes) — like a Deployment, but for workloads that need a stable identity: each replica gets a fixed ordinal name (db-0, db-1, db-2) and its own PersistentVolumeClaim that follows it across rescheduling, and scale-up/down happens strictly in order — the right controller for databases and other clustered stateful software, where Deployment's interchangeable, identity-less pods would be actively wrong. (Ch. 6)
  • STS / AssumeRole — AWS Security Token Service issues short-lived, temporary credentials when a principal "assumes" an IAM role, which is the mechanism behind cross-account access, EC2/EKS instance roles, and OIDC-based CI credentials (above) — none of which require a long-lived access key ever to exist. (Ch. 5)
  • Subscription (Azure) — Azure's billing-and-isolation boundary, the direct analog of an AWS account: every resource lives in exactly one subscription, quotas (GPU family quotas among them) are tracked per subscription, and a subscription is the smallest unit Azure ever bills separately. Sits below a management group (above) in the ARM hierarchy. (Ch. 16g)
  • Supernode (graph) — a node whose degree sits orders of magnitude above a graph's average, breaking every traversal cost model that assumes roughly constant fan-out per hop: a single supernode on a path replaces that hop's expected fan-out with its own, much larger one, and a query planner estimating from the graph's average degree never sees it coming. Common causes are a shared value (an address, a jurisdiction) promoted to a node with no size cap, or an over-merge in entity resolution (above) pooling several distinct entities' relationships onto one ID. The standing mitigation is treating the degree distribution as a health check, not a one-off: flag and cap traversal through any node far above the mean. (Ch. 19d)
  • Supervisor multi-agent — an orchestration topology where a supervisor agent classifies an incoming task and routes it to one of several sub-agents, each holding its own curated, narrow tool catalog rather than every sub-agent (or one monolithic agent) seeing the organization's full tool catalog. Keeps each sub-agent's fixed per-call overhead \(b\) small and its tool-selection accuracy high — the direct fix for the forty-tool problem's 96%-to-81% accuracy drop. (Ch. 19e)
  • SVD (Singular Value Decomposition) — the matrix factorization \(X = U\Sigma V^\top\) that underlies PCA (above): \(U\)'s columns are the left singular vectors, \(V\)'s columns (the principal-component directions) are the right singular vectors, and \(\Sigma\)'s diagonal holds the singular values, whose squares are proportional to the variance explained by each component — computing PCA via SVD directly on the (centered) data matrix is more numerically stable than explicitly forming and eigen-decomposing the covariance matrix. Full derivation in Appendix D. (Ch. 17)
  • Taints / tolerations (Kubernetes) — a taint on a node repels pods from scheduling there unless the pod carries a matching toleration — the mechanism behind dedicating nodes to a specific workload (GPU nodes tainted so only GPU-requesting pods land there) without needing an allow-list on every other pod in the cluster. (Ch. 6)
  • Terminal condition (agent) — any condition under which an agent loop must stop transitioning and report, rather than looping again: a final answer, an exhausted step/token/dollar/wall-clock budget, an unrecoverable tool error, or a human reviewer's rejection at a validation gate. Coding all four explicitly, with a structured reason string, is what lets a stuck model, a budget blowout, and a rejected action be told apart from the outside instead of collapsing into one undifferentiated "it stopped." (Ch. 19e)
  • Tensor core — a GPU execution unit specialized for small matrix multiply-accumulate (MMA) operations (\(D = A \times B + C\) over a tile), as opposed to a conventional CUDA core's one scalar FMA per instruction per thread. Tensor cores are why lower-precision formats (tf32, bf16, fp8) deliver dramatically higher peak throughput than fp32 on the same silicon — the accumulator is kept at higher precision than the operands on every current tensor-core path, which is what keeps a matmul's summed error bounded even at low operand precision. (Ch. 16a)
  • Terraform module / workspace — a module is a reusable, parameterized bundle of .tf resource blocks (inputs as variable, outputs as output) called from elsewhere with module "name" { source = ...; ... }, the standard way to avoid copy-pasting the same VPC/EKS-cluster/RDS boilerplate across environments. A workspace is a named, isolated instance of Terraform state (above) for the same configuration — terraform workspace new staging gives you a separate state file so dev/staging/prod don't collide, without duplicating the .tf code itself. Workspaces are a lighter-weight alternative to fully separate state backends per environment; many teams prefer separate backends (or separate root modules) for environments with meaningfully different blast radius, reserving workspaces for short-lived, throwaway variants (a per-developer sandbox, a per-PR preview environment). (Ch. 7)
  • tf.function / retracing — the TensorFlow decorator that traces a Python function into a static graph the first time it runs with a given input signature (shapes + dtypes), then replays that compiled graph on subsequent calls instead of re-executing Python. A retrace happens whenever a new signature is seen — a different shape, a Python scalar changing value where a tensor was expected — and is silent by default: no exception, just a slower call while a new graph compiles. A serving endpoint fed unbucketed, variable-length inputs can retrace on nearly every request, which reads as a p99 latency spike, not a logged error. (Ch. 16c)
  • Time-travel debugging (agent) — loading a checkpoint from an earlier step of a completed (often failed) agent trajectory and re-running forward from there, deterministically, to inspect exactly what the model decided and why — the agent-loop analogue of resuming a debugger at a breakpoint, made possible only because deterministic replay (above) reads recorded observations instead of re-calling live tools. (Ch. 19e)
  • TLS / X.509 / CA — Transport Layer Security encrypts and authenticates a connection; an X.509 certificate binds a public key to an identity (a domain name) and is itself signed by a Certificate Authority, forming a chain of trust up to a root CA that clients trust by default; mTLS (mutual TLS, below under Zero trust) additionally requires the client to present a certificate too. (Ch. 2, Ch. 9)
  • Tokenization — splitting raw text into the discrete units (tokens) a language model actually consumes; modern LLMs use subword tokenization (BPE, above, or similar) rather than whole words, which bounds vocabulary size while still being able to represent any string, at the cost of common words costing 1 token and rare/foreign words costing several. (Ch. 19)
  • Tombstone (ingestion sense) — a soft-delete marker (status = 'tombstoned', tombstoned_at = now()) set on a row that stays physically present, excluded from every active-only query, and physically removed only later by a scheduled job once a grace window passes — never an immediate DELETE, so a document briefly invisible at the source due to a transient outage can still be restored before phase two runs. This is a different mechanism from the CDC tombstone (a null-valued message published to a compacted log to trigger stream compaction): the two share a name and an intuition — "mark gone, don't erase" — at two different layers, wire format versus row-level status flag. (Ch. 19b)
  • Tombstone (HNSW graph sense) — a third, distinct meaning of the same word: deleting a node from an HNSW graph (above) marks it dead in a bitmap rather than severing its edges, because a naive removal can disconnect the graph if the node bridged two otherwise-distant regions. A query's greedy descent still traverses through a tombstoned node's edges (preserving navigability) but the node itself is never returned. Tombstoned nodes cost exactly as much memory and traversal time as live ones, so a growing tombstoned fraction silently degrades recall at a fixed efSearch (above) even though no HNSW parameter changed; the only fix is compaction, a scheduled rebuild that drops tombstoned nodes and re-links their orphaned neighbors. (Ch. 19c)
  • Tool versioning (agent) — an additive-only discipline for changing an agent tool's schema: new fields are optional with a documented default, existing field names/types/semantics never change underneath a version number, and breaking changes ship as a new tool name rather than a mutation of the old one. A non-additive change can invalidate every checkpointed trajectory that has not resumed yet, since a resumed session replays old tool-call arguments against a validator that no longer accepts their shape. (Ch. 19e)
  • torch.compile (TorchDynamo / Inductor) — PyTorch's just-in-time compiler: TorchDynamo intercepts Python bytecode as a function runs and captures the operations into an FX graph (falling back to eager execution, not erroring, on constructs it cannot trace); Inductor lowers that graph to fused, often faster kernels (Triton on GPU). The captured graph is guarded by the exact input shapes/dtypes/values Dynamo saw — a guard failure (a new shape, a changed control-flow branch) triggers a silent recompilation, not an exception, which is why a training loop's step time can be flat for thousands of steps and then start spiking with no error in the logs once inputs start varying in shape. (Ch. 16c)
  • Trajectory (agent evaluation) — the complete, ordered sequence of (state, action, observation) triples an agent loop produces from start to a terminal condition (above). Trajectory-level evaluation — scoring the whole sequence, including which tools were called and in what order — catches a wrong path to a right final answer that a final-answer-only proxy silently rewards, and is the measurement cost-per-task-resolved (above) depends on for a trustworthy success rate. (Ch. 19e)
  • Transfer learning — reusing a model already trained on a large, general dataset (a "pretrained" model) as the starting point for a new, usually smaller, more specific task, instead of training from randomly initialized weights. The pretrained model's early/middle layers have already learned broadly useful representations (edges and textures for vision, syntax and world knowledge for language), so fine-tuning (above) only needs to adapt the last layers (or a low-rank update, as in LoRA, below) to the new task's specifics — dramatically reducing both the labeled data and the compute a new task requires versus training from scratch. Nearly every practical deep-learning system in this book (a fine-tuned LLM, a ResNet backbone reused for a custom detector) is an instance of transfer learning. (Ch. 18, Ch. 19, Ch. 20)
  • Transformer — the attention-based (above) sequence architecture that replaced recurrence (RNNs/LSTMs) as the default for text, and increasingly for images (ViT, above) and other modalities: stacked blocks of self-attention (mix information across positions) plus a position-wise feed-forward network (transform each position independently), with residual connections (above) and layer normalization around each sub-layer. Its key practical advantage over RNNs is that all positions can be processed in parallel during training (no sequential recurrence to unroll). (Ch. 19)
  • Trust boundary (LLM-specific) — the point past which content can no longer be distinguished from instructions once it is inside an LLM's context window. A conventional application enforces this boundary at a well-defined interface the untrusted content never influences (a SQL driver's parameter binding, a shell's argv array); a transformer has no equivalent, because system prompt, retrieved document, tool output, and user message are concatenated into one token sequence and mixed by the same attention weights with no bit marking which span came from where. Every LLM security control is therefore placed before that concatenation or entirely outside the model, never inside the token sequence itself. (Ch. 19f)
  • TTFT / TPOT — the two LLM-serving latencies a user actually experiences. TTFT (time to first token) is the wait from request arrival to the first output token, dominated by prefill's compute-bound pass: \(\text{TTFT} \approx 2N\cdot\text{tokens\_in}/P_{\text{peak}}\). TPOT (time per output token) is the steady-state gap between subsequent tokens during decode, dominated by the memory-bandwidth-bound read of weights plus KV cache (above) every step. Worked example: an 8B-parameter model, 2,000-token prompt, \(P_{\text{peak}}=3.12\times10^{14}\) FLOP/s: \(\text{TTFT} \approx 2\times8\times10^9\times2{,}000 / 3.12\times10^{14} \approx 103\) ms. A short prompt with a long generation is TPOT-dominated (feels "slow to keep typing"); a long prompt with a short answer is TTFT-dominated (feels "slow to start"). (Ch. 16e)
  • Two-column disaster — the failure mode where naive, stream-order PDF text extraction interleaves two (or more) columns of a page's actual reading order into one incoherent string, because a PDF's content stream encodes drawing order, not reading order, and a multi-column layout is exactly where those two orders diverge. The extraction does not fail — it returns a plausible, grammatically-almost-coherent string with no exception — which is what makes it dangerous: a monitoring dashboard watching for parse errors sees a healthy pipeline while every chunk built from that page splices two unrelated sentences. The mitigation is column-aware extraction: cluster text spans into columns by their x-coordinates and emit each column fully before moving to the next. (Ch. 19b)
  • VAE (Variational Autoencoder) — a generative autoencoder (above) whose encoder outputs the parameters of a distribution (mean and variance of a Gaussian) over the latent code rather than a single point, trained by maximizing the ELBO (above); because nearby points in its latent space decode to similar, plausible outputs, you can sample new data by sampling latent points from the prior and decoding. (Ch. 21)
  • Vanishing / exploding gradient — the failure mode where backpropagation (above) through many layers multiplies many partial derivatives together, and if those derivatives are consistently \(<1\) in magnitude the product shrinks toward zero (vanishing — early layers barely update, training stalls), while if they are consistently \(>1\) it grows without bound (exploding — weights blow up to NaN). Worked example: the sigmoid activation's derivative \(\sigma'(x)=\sigma(x)(1-\sigma(x))\) has a maximum value of exactly \(0.25\) (at \(x=0\)); through just 10 sigmoid layers, the best-case gradient magnitude is multiplied by \(0.25^{10} \approx 9.5\times10^{-7}\) — effectively zero, which is precisely why deep sigmoid networks were nearly untrainable before ReLU (above), residual connections (above), and batch normalization (above) became standard: ReLU's derivative is exactly 1 for any positive input (no shrinkage), residual connections give the gradient a +x path that bypasses the multiplicative chain entirely, and gradient clipping (above) caps the exploding side directly. (Ch. 18, Appendix D)
  • Vectorization — expressing a computation as array-level operations (NumPy/PyTorch ufuncs, matrix multiplies) executed by compiled, often SIMD- or GPU-parallelized code, instead of an explicit Python for loop over elements; the same logical computation can run 10–1000× faster purely from avoiding per-element Python interpreter overhead, with no change in what is mathematically computed. (Ch. 11)
  • VPC (Virtual Private Cloud) — a private, software-defined network within an AWS Region: you define its CIDR range (above), subnets (public — has a route to an internet gateway; private — doesn't), route tables, and gateways; almost everything else in AWS (EC2, RDS, EKS nodes) is placed inside a VPC's subnets. (Ch. 5)
  • Walk-forward validation / backtesting — the time-series-correct alternative to shuffled cross- validation: train on data up to time \(t\), evaluate on the next window after \(t\), then slide the whole window forward and repeat — never evaluating on data that precedes the training window chronologically. A backtest is the broader practice of simulating a trading (or any sequential-decision) strategy over historical data as if it had been run live, and is only trustworthy when it respects this same no-peeking-into-the-future discipline (see also CPCV, above). (Ch. 23)
  • Warp (GPU) — the real unit of execution on an NVIDIA GPU: a fixed group of 32 consecutive threads from a block, scheduled and executed together in lockstep (SIMT — Single Instruction, Multiple Threads). A block's threads are launched in whatever grouping the programmer chose, but the hardware always schedules them 32 at a time; when threads within a warp take different branches (warp divergence), the hardware serially executes each path with the non-participating lanes masked off, so divergent branches cost the sum of both paths' time, not the max. (Ch. 16a)
  • Watermark (stream processing) — a streaming engine's declared bound on "how late can an event still arrive and be counted," e.g. "no event older than 10 minutes behind the current processing time will be included in a window's aggregate"; it's the mechanism that lets a streaming aggregation ever close a time window and emit a final result, at the cost of silently dropping (or routing to a "late data" side-output) anything arriving after the watermark has passed. (Ch. 12)
  • WebDataset shard — a tar-formatted archive bundling many training samples (image bytes, label, any auxiliary metadata, grouped per-sample) into one sequentially-readable file, typically 100 MB to a few GB each, so reading a dataset means streaming through a small number of large files rather than opening one file per sample. The structural fix for the metadata wall (above): \(N_f\) collapses from one-per-sample to one-per-shard, often three or four orders of magnitude smaller for the same dataset — \(10^6\) files repacked as 1,000 shards cuts open-overhead from 312.5 s to about 0.3 s. The trade: true per-sample random access becomes expensive again, usually answered with a bounded in-memory shuffle buffer that shuffles within and across a rolling window of shards rather than the whole dataset at once. (Ch. 16b)
  • WEKA — a software-defined, NVMe-first parallel filesystem, sold as a managed or appliance product across cloud and on-prem deployments, architected from the ground up around all-flash media and low metadata latency rather than adapted from disk-era assumptions the way Lustre (above) and GPFS (above) originally were. Exposes POSIX, S3, and native GPUDirect Storage (above) support; typically reached for when metadata-heavy, latency-sensitive access patterns (many small files, heavy random I/O) matter as much as or more than raw sequential throughput. (Ch. 16b)
  • Weight decay — an optimizer-level implementation of L2 regularization (above) that shrinks every weight by a small multiplicative factor at each step, independently of the gradient: \(\theta \leftarrow (1-\lambda\eta)\theta - \eta\nabla_\theta L\). In plain SGD this is mathematically identical to adding \(\lambda\lVert\theta\rVert_2^2\) to the loss, but for adaptive optimizers like Adam (above) the two are not equivalent — Adam's per-parameter adaptive step size interacts badly with an L2 penalty folded into the gradient, which is why AdamW (Adam with decoupled weight decay, applying the shrinkage directly to the weights outside the adaptive-gradient computation) is the standard optimizer for training Transformers rather than plain Adam with an L2 term. (Ch. 18, Ch. 19)
  • Well-Architected Framework (AWS) — AWS's five-pillar design checklist for reviewing an architecture: Operational Excellence, Security, Reliability, Performance Efficiency, and Cost Optimization (a sixth pillar, Sustainability, was added later); used less as a rulebook and more as a structured set of questions to ask about any nontrivial design before it ships. (Ch. 5)
  • Working memory (agent) — an agent loop's own context window: system prompt, tool schemas, and the current trajectory's accumulated history — the thing that grows quadratically with trajectory length and that compaction (above) exists to bound. Distinct from the other three agent memories in that it holds only the current trajectory and disappears once that trajectory ends, unless something in it was written out to episodic or semantic memory first. (Ch. 19e)
  • ZeRO (Zero Redundancy Optimizer) — a family of three sharding stages that eliminate data parallelism's per-GPU state redundancy without changing the model itself. ZeRO-1 shards only the optimizer state; ZeRO-2 additionally shards gradients (both add no extra communication over plain DDP, above — same reduce-scatter-then-all-gather total bytes). ZeRO-3 additionally shards the parameters themselves, adding a per-layer all-gather in both forward and backward, for roughly \(1.5\times\) plain DDP's wire volume in exchange for memory that no longer scales with model size per GPU. FSDP (above) is PyTorch's native ZeRO-3 implementation. (Ch. 16d)
  • Zero-shot / few-shot / in-context learning — three ways to get an LLM to perform a task without any gradient update to its weights, purely through the prompt (contrast with fine-tuning, above, which does update weights). Zero-shot gives only an instruction ("Classify this review as positive or negative"). Few-shot additionally includes a handful of worked examples directly in the prompt before the real query — the model conditions its next-token predictions on those examples' pattern without any training step, purely by the Transformer's attention (above) over the prompt's tokens; this ability to pick up a task's pattern from a few in-prompt examples alone is called in-context learning, and it emerges as a capability mainly in models above a certain scale rather than being explicitly programmed. Few-shot prompting typically improves accuracy over zero-shot at the direct cost of a longer prompt (more tokens billed and a larger share of the context window consumed before the actual query even appears). (Ch. 19)
  • Zero trust / mTLS — a security model that assumes no implicit trust from network location alone (being "inside the VPC" grants nothing by itself); every request is authenticated and authorized on its own merits, commonly via mutual TLS (mTLS) — both sides of a connection present and verify a certificate — a pattern that service meshes (Istio, Linkerd) automate transparently between pods. (Ch. 9)

Cross-cutting maps

Two relationships recur so often across chapters that they deserve a diagram of their own rather than being scattered across a dozen separate bullets above.

The training loop, which every entry from Loss function to Optimizer to Backpropagation is a piece of:

flowchart LR
    A["Data batch"] --> B["Forward pass<br/>(model prediction)"]
    B --> C["Loss function<br/>(cross-entropy, MSE...)"]
    C --> D["Backpropagation<br/>(∂L/∂θ via chain rule)"]
    D --> E["Optimizer step<br/>(SGD / Adam)"]
    E --> F["Updated weights θ"]
    F -.next batch.-> A

The identity/access triangle, which every entry from IAM to OIDC to RBAC to ESO is a piece of:

flowchart TD
    Human["Human user"] -->|"OIDC login (Keycloak)"| Proxy["oauth2-proxy / Ingress"]
    CI["CI/CD pipeline"] -->|"OIDC token exchange"| STS["AWS STS AssumeRoleWithWebIdentity"]
    STS --> Temp["Short-lived AWS credentials"]
    Pod["Kubernetes pod"] -->|"Service Account token"| RBAC["Kubernetes RBAC"]
    Pod -->|"synced by"| ESO["External Secrets Operator"]
    ESO -->|"reads from"| SM["AWS Secrets Manager"]
    Temp --> IAMPolicy["IAM Policy (least privilege)"]
    RBAC --> IAMPolicy

Both diagrams point at the same underlying lesson: every layer of the stack re-implements the same two patterns — a reconciliation loop (desired state vs. actual state, whether that's Kubernetes' controllers or gradient descent driving a loss to zero) and a chain-of-trust for identity (whether that's an X.509 certificate chain or an IAM AssumeRole chain). Recognizing the pattern is more valuable than memorizing each instance of it.

Pitfalls & best practices

Same word, different meaning across teams

state, policy, deployment, and pipeline each mean something different depending on whether the speaker means it in the Terraform/infra sense or the Kubernetes/RL/CI sense. In a cross-functional design review, the single highest-leverage clarifying question is: "which <term> do you mean — the infra one or the ML one?" This appendix's per-domain sub-definitions (e.g. Kernel, Policy, Namespace) exist specifically to make that ambiguity visible rather than silently assumed away.

  • Don't let a glossary go stale. A term whose definition no longer matches how the codebase actually uses it is worse than no definition — it actively misleads. Treat this file like any other doc: update it in the same pull request that introduces or repurposes a term, not "later."
  • Precision over brevity when it's load-bearing. "Roughly the loss" is not a definition — a reviewer who reads "~cross-entropy~" without the formula cannot verify whether your code computes it correctly. Every entry above that has a formula also has the formula, not a hand-wave.
  • Don't compute a metric you can't also compute by hand once. If you cannot reproduce a Sharpe ratio, an IoU, or a KL divergence on a 3-number toy example with pencil and paper, you cannot sanity-check your code's output against a bug — you can only trust it blindly. Every worked example above is small enough to redo by hand in under a minute; do that at least once for any metric you rely on operationally.
  • Multiple-comparisons blindness. Reporting the best of \(N\) tuned models/strategies without correcting for \(N\) (Deflated Sharpe Ratio, above, is the finance-specific instance; the same logic applies to hyperparameter search leaderboards and A/B test dashboards with many simultaneous variants) systematically overstates how good the winner really is.
  • Confusing liveness with readiness, confusing a Security Group with a NACL's statefulness, or patching a creationPolicy: Owner ESO-managed Secret by hand are three of the most common operational glossary-adjacent mistakes — each looks like it works immediately and fails silently later (a restart loop, an unexpectedly open port, a secret key that vanishes on the next sync).

Exercises

  1. Little's law, inverted. A batch-inference service must sustain \(\lambda = 200\) requests/s. Your SLO requires an average of at most \(L=20\) requests in flight at any time (to bound memory). What is the maximum average time-in-system \(W\) you can afford? What does this imply about how fast each request's model call must complete?
  2. Cross-entropy by hand. A 4-class classifier outputs \(\hat p = [0.05, 0.15, 0.70, 0.10]\) for an example whose true class is the 4th. Compute the cross-entropy loss in nats, then recompute it assuming the model had instead put \(\hat p_4 = 0.40\) with the other three probabilities rescaled proportionally. By how much did the loss change, and why is the relationship nonlinear?
  3. KL divergence asymmetry. Using \(p=[0.9,0.1]\) and \(q=[0.5,0.5]\), compute both \(D_{\mathrm{KL}}(p\|q)\) and \(D_{\mathrm{KL}}(q\|p)\). Confirm they are not equal, and explain in one sentence why KL divergence is not a valid distance metric.
  4. IoU threshold sensitivity. Two bounding boxes are \(A=[0,0,20,20]\) and \(B=[10,10,25,25]\). Compute their IoU. Would this pair count as a match under the common \(\text{IoU} \ge 0.5\) detection threshold? At what threshold does it stop counting as a match?
  5. Deflated Sharpe intuition. You backtested \(N=100\) independent strategy variants and the best one shows an annualized Sharpe of 1.5, with per-trial Sharpe standard error \(\sigma_{SR}=1\). Using the approximation \(\mathbb E[\max \widehat{SR}] \approx \sigma_{SR}\sqrt{2\ln N}\), estimate the Sharpe you'd expect from pure luck alone with 100 trials. Is 1.5 still impressive? What would you need to do differently next time (see Ch. 23) to make a defensible claim of real edge?
  6. Token bucket capacity. A rate limiter uses a token bucket with burst capacity \(B=50\) and refill rate \(r=5\) tokens/s. A client sends a burst of 50 requests instantly (bucket starts full), then keeps sending at a steady 8 requests/s. After the initial burst, how long until the client starts seeing 429s, and at what steady rate does it then get accepted vs. rejected once the bucket reaches equilibrium?
  7. Cosine annealing by hand. Using \(\eta_{\max}=2\text{e-}4\), \(\eta_{\min}=0\), and \(T=1000\) total steps, compute the learning rate \(\eta_t\) at \(t=250\), \(t=500\), and \(t=750\). Sketch (in words) why the rate falls faster in the middle of training than near the very start or very end.
  8. One-page cross-functional map. Pick any two chapters that are non-adjacent in the book's arc (e.g. Terraform and Reinforcement learning). Find one term from this glossary that appears in both, in a different sense each time (hint: state, policy), and write two or three sentences explaining the two meanings to someone who only knows one of the two chapters.