This article was generated entirely by AI without human authorship. Please read with discretion, or have an AI verify its accuracy.
AI TranslationSimplified ChineseEnglish
Summary:
The green book is pretty much required reading if you're getting into AI inference — once you've read it, you have a basic map of what this field is about. I've organized the terms here for easy reference later.
Inference Engineering: A Glossary of Inference Engineering Terms, Organized by System Pipeline
Once a generative AI app actually ships, the questions stop being just "can the model answer?" and start including how the model gets loaded, how requests are queued, how the GPU computes and moves data, and how the service stays stable under real traffic.
This article is compiled from Inference Engineering by Philip Kiely. The book is published by Baseten Books, with an official online reading version. The terms below come mainly from the book's Appendix A: Inference Glossary. The book lists them alphabetically; here I've regrouped them by inference system pipeline so you can cross-reference them while writing code, reading source, or looking at performance metrics.
The book spans everything from CUDA, GPUs, and inference engines to KV cache, quantization, parallelism, diffusion models, speech, multi-cloud deployment, and production serving. I got through it in two days — and Appendix B is a great indexing resource too.
1. Models, Applications, and Data Representation
This group answers "what a model is, how an application calls it, and how text becomes a representation the model can process."
Term
What it means
What to look at in an inference system
Activation function
Activation function: a mostly differentiable nonlinear function inserted between linear layers, e.g. ReLU.
Without nonlinearity, a multilayer network collapses into a single matrix multiplication.
Generative AI
Generative AI: learns patterns from data and generates new text, images, audio, video, or code.
The emphasis is on "generating new content," as opposed to traditional ML, which only classifies or predicts.
Machine learning (ML)
Machine learning: learns predictive models from data, e.g. classification and trend forecasting.
In the glossary, it serves as the contrast to generative AI.
Foundation model
Foundation model: trained on broad data and usable as the base for many downstream tasks.
You can prompt it directly, or fine-tune it into a domain-specific model.
Open model
Open model: model weights are freely available, e.g. Llama, DeepSeek, Whisper.
Visible weights usually mean more control over deployment, quantization, and hardware choices.
Closed model
Closed model: proprietary model whose weights are not available, e.g. GPT, Claude, Gemini.
Typically used through an API — what you control is requests, routing, and cost, not the underlying weights.
Transformer
Transformer: the foundational network architecture behind generative AI.
Attention, Q/K/V, KV cache, and most inference optimizations revolve around it.
Causal language model (CLM)
Causal language model: a decoder-only Transformer that predicts the next token using only the preceding context.
Autoregressive LLM inference is essentially running a CLM's next-token prediction loop.
Large Language Model (LLM)
Large language model: takes a text prompt and generates a new sequence of text.
Typical families include GPT, Claude, Llama, and DeepSeek.
LLM
Short for `Large Language Model`.
Same concept as the row above; in engineering docs it's usually written simply as LLM.
Generative Pretrained Transformer (GPT)
Generative Pretrained Transformer: the family of text-generation LLMs created by OpenAI.
It's the name of a model family, not a catch-all for all LLMs.
Agent
Agent: an AI application that doesn't just answer questions but also calls tools and takes action.
A single user request may trigger multiple inference calls, multiple models, and multiple modalities.
AI-native application
AI-native application: a product whose core experience and value depend on generative models.
Modality, latency budget, unit economics, and usage patterns all feed back into the inference architecture.
Application Programming Interface (API)
API: a structured interface for sending requests and receiving responses.
Inference engines typically expose model querying through an API.
Inference
Inference: serving an AI model in production.
The emphasis here is on serving, not just "running a single forward pass."
Inference engine
Inference engine: a high-performance runtime supporting optimizations such as batching, caching, quantization, and speculation.
vLLM, SGLang, and TensorRT-LLM all fall into this category.
Prompt
Prompt: the instruction given to the model; for diffusion models it may also include a negative prompt, step count, and guidance parameters.
Prompt length directly affects prefill, TTFT, and KV cache usage.
Chat template
Chat template: serializes roles, delimiters, and sequence start/end tokens into input as the model requires.
Change the template for the same conversation and the actual token sequence may differ.
Token
Token: the basic unit of text an LLM processes — essentially an integer representing a string fragment.
Latency, throughput, context window, and billing are usually measured in tokens.
Tokenizer
Tokenizer: performs deterministic conversion between strings and token sequences.
Different models use different tokenizers; a more efficient tokenizer can lower end-to-end latency.
Vocabulary
Vocabulary: the full set of tokens a model uses to represent data.
Vocabulary size affects the logits vector and tokenization behavior.
Input sequence
Input sequence: the tokens in a request handed to the model, processed during prefill.
The longer the input, the higher the cost of prefill computation and KV cache setup.
Input Sequence Length (ISL)
Input sequence length: the number of input tokens in a single request.
Together with OSL, it describes the length shape of a workload.
Output sequence
Output sequence: the tokens the model generates during decode.
Output length determines how long decode runs and what the user sees streamed.
Output Sequence Length (OSL)
Output sequence length: the number of output tokens generated by a single request.
A long OSL tends to amplify decode, KV cache, and queue pressure.
Context window
Context window: the total cap on input, reasoning, and output tokens a model can process in a single request.
It's both the ceiling on model capability and a constraint on VRAM, KV cache, and scheduling.
Function calling
Function calling, also known as tool calling / tool use: the model picks a function from a given set and returns structured arguments.
The server must validate the schema, execute the tool, and feed the result back to the model.
Structured output
Structured output: model output that follows a specified schema.
The book stresses achieving this with generation constraints such as logit bias, not just by agreeing on a prompt.
Logit biasing
Logit biasing: adjusting or constraining token probabilities before sampling to steer output such as JSON or tool calls.
It sits after logits are produced and before the final sampling step.
2. Request Lifecycle, Decoding, and Performance Metrics
These terms cover "what happens after a request comes in, and how should speed be measured."
Term
What it means
What to look at in an inference system
Autoregressive token generation
Autoregressive token generation: each new token depends on the tokens already generated.
This naturally forms a step-by-step decode loop and is the backdrop for KV cache and speculative decoding.
Pretraining
Pretraining: large-scale training on broad corpora to produce a foundation model.
It happens before serving; inference engineering typically consumes the weights it produces.
Training
Training: learning model weights from data through backpropagation and optimization.
Training is compute-heavy and usually relies on large-scale GPU clusters.
Prefill
Prefill: the LLM inference phase that processes the input sequence in one pass and builds the KV cache.
It is usually compute-bound, and a long ISL can significantly increase TTFT.
Decode
Decode: the phase in which the autoregressive loop generates one token at a time.
This phase tends to be constrained by memory bandwidth and KV reads; each forward pass usually adds just one token.
Sampling (decode)
Sampling: selecting the next output token from the generated logits.
Common strategies include greedy/argmax, temperature, top-k, and top-p.
Logits
Logits: the vector of unnormalized scores the model produces for every token in the vocabulary on each forward pass.
Sampling, temperature, and logit bias all operate on it or on a transformation of it.
Softmax
Softmax: turns attention scores into probabilities, and also normalizes the logits during decode into a probability distribution.
It bridges the two stages of scoring and probability-based sampling.
Temperature
Temperature: controls how random token selection is; low values are more deterministic, high values more diverse.
It changes the sampling distribution; it does not change the model's underlying capabilities.
Batch
Batching: processing multiple inputs at once.
It improves device utilization, but differences in request length lead to waiting or padding.
Batch sizing
Batch size: a core lever for trading off latency against throughput.
A larger batch usually raises total throughput, but it worsens latency for each individual user.
Dynamic batching
Dynamic batching: fires once the batch is full or a short timer expires.
It balances utilization and latency stability; for LLMs it has largely been replaced by continuous batching.
Continuous batching (in-flight)
Continuous batching / in-flight batching: interleaves requests at the token level so GPU slots always have work to do.
This matters especially for online LLM requests that vary in length and arrive continuously.
In-flight batching
Another name for `continuous batching`.
The key point is that requests can join, finish, and exit while a batch is still executing.
Chunked prefill
Chunked prefill: splitting a long input into chunks and overlapping them with decode or other work.
It keeps one very long sequence from monopolizing resources for an extended period.
Queue (request)
Request queue: holds traffic that exceeds current capacity until autoscaling spins up new replicas.
Queue length and wait time are signals of insufficient capacity, cold starts, or traffic spikes.
Online inference
Online inference: serving requests in real time, prioritizing strict latency budgets.
Watch TTFT, ITL, P99, and availability.
Offline inference
Offline inference: processing large jobs asynchronously in batches, prioritizing throughput and cost.
You can trade per-request latency for higher device utilization.
Local (edge) inference
Local / edge inference: running inference on end-user devices such as phones and laptops.
It cuts out network round trips, but is limited by on-device compute, memory, and power.
Latency percentiles
Latency percentiles: using P50, P90, P95, P99, and so on to see the distribution behind average-to-tail latency.
P99 is often closer to the worst user experience than the average is.
Time to first byte (TTFB)
Time to first byte: the time from request start to receiving the first output byte.
It is shaped by the network, server-side processing, and the streaming protocol together.
Time to first token (TTFT)
Time to first token: the time from request start to receiving the first output token.
This is critical for interactive scenarios like chat and code completion, and it usually includes prefill.
Inter-token latency (ITL)
Inter-token latency: the interval between consecutive generated tokens during decode.
The smaller the ITL, the smoother the streamed output feels to users; 2 ms, for example, corresponds to roughly 500 TPS.
Perceived TPS
Perceived TPS: the number of tokens per second a single user actually sees in the streamed output.
It reflects the individual user experience more closely than aggregate throughput does.
Tokens per second (TPS)
Tokens per second: the number of tokens streamed to the end user each second; in the glossary this refers to perceived TPS.
Be clear about whether you mean single-user TPS or aggregate throughput.
Throughput
Throughput: the total work completed per unit of time, such as total tokens/s.
Higher aggregate throughput does not mean better TPS or TTFT for each user.
Baselines
Baselines: carefully recorded performance and quality measurements taken before optimization.
Without a baseline, you cannot reliably attribute gains or regressions to an optimization.
Load testing
Load testing: sending sustained high traffic to probe throughput limits, queue behavior, and scaling.
Pair it with realistic ISL/OSL distributions rather than uniform synthetic requests alone.
Jitter traffic (bench)
Jitter traffic: adding randomness to request arrival times and sequence shapes.
This comes closer to real workloads than tidy, fixed-interval traffic.
Real-time factor (RTF)
Real-time factor: a speed metric for ASR transcription; if one hour of audio finishes transcribing in six seconds, the RTF is 600.
It is well suited to expressing audio processing speed relative to the original playback duration.
3. Attention, the KV Cache, and Context Optimization
Term
What it means
What to look at in an inference system
Attention
Attention: the Transformer mechanism that links the current token to historical tokens through Q/K/V projections and softmax.
It is heavy in both compute and memory traffic, making it a prime target for inference optimization.
Head (attention)
Attention head: a single independent attention computation within a layer.
The multi-head structure lets the model capture relationships from different subspaces.
Cross-attention
Cross-attention: the Q of one sequence is conditioned on the K/V of another sequence.
It shows up often in text-conditioned image generation, multimodal models, and diffusion denoising pipelines.
KV cache
KV cache: stores the K/V tensors for each historical token so history projections are not recomputed at every step.
It trades VRAM for compute; the longer the context, the larger the cache and the more there is to read at each step.
Prefix caching
Prefix caching: reusing the KV of shared prefixes across requests to skip redundant prefill.
It can significantly improve TTFT for code completion, multi-turn conversations, and agent systems.
Cache-aware routing
Cache-aware routing: sending requests to replicas that already hold the matching prefix or the required LoRA.
Routing considers cache hits and local state, not just load.
PagedAttention
PagedAttention: storing KV blocks in fixed-size pages to manage long-context memory more effectively.
Its typical benefits are less fragmentation and higher KV utilization; do not conflate it with ordinary OS paging.
Rotary positional embeddings (RoPE)
Rotary positional embeddings: representing position with learnable rotations, which helps with long-context extrapolation.
There is a trade-off between long-context capability and attention memory requirements.
Context Parallelism (CP)
Context parallelism: replicating weights across GPUs while splitting the attention context.
It suits models with extremely large contexts or video latent spaces.
Ring attention
Ring attention: a CP mechanism in which GPUs pass partial attention results around in a ring.
It enables longer contexts by reducing all-to-all pressure.
Disaggregation
Disaggregated inference: splitting prefill and decode into engines that scale independently and use different hardware resources.
It lets two very different bottlenecks scale separately, but adds transfer and scheduling complexity.
4. GPUs, the Memory Hierarchy, and the CUDA Software Stack
This group of terms explains where the model actually computes, how data moves, and how kernels execute.
Term
What it means
What to look at in an inference system
Central Processing Unit (CPU)
CPU: a general-purpose processor suited to sequential workloads.
Mostly handles orchestration, scheduling, networking, and preprocessing; it usually doesn't do the main generation compute itself.
Graphics Processing Unit (GPU)
GPU: a highly parallel processor, now widely used for both training and inference in generative AI.
Look at memory capacity, bandwidth, compute units, and interconnect—not just the model name.
GPU node
GPU node: typically a standard chassis holding 8 GPUs connected via NVLink/NVSwitch.
It's the common building block for multi-GPU inference and intra-node parallelism.
Node
Node: a physical 8-GPU base unit with NVLink/NVSwitch; multiple nodes are then linked with InfiniBand.
Clarifies the boundary between intra-node and inter-node communication.
Instance (cloud)
Cloud instance: a configured virtual machine that includes GPU, CPU, RAM, storage, networking, and interconnect.
On the cloud, billing, deployment, and autoscaling usually happen at the instance level.
High-Bandwidth Memory (HBM)
High-bandwidth memory: the high-bandwidth memory used as VRAM in data center GPUs, such as HBM3, HBM3e, and HBM4.
It determines how much room weights, activations, and KV have, and how fast data can be fed in.
VRAM (device memory)
VRAM: the device memory on the GPU that holds weights, KV, and activations.
Capacity limits model size and KV headroom; bandwidth affects decode TPS.
Out-of-memory error (OOM)
Out-of-memory error: the GPU doesn't have enough VRAM to load weights or run inference.
Troubleshoot model precision, context length, batch, KV cache, and parallelism together.
Bandwidth
Bandwidth: how much data a memory or interconnect can move per second.
Decode is often bandwidth-bound; NVLink, HBM, and InfiniBand are bandwidth at different levels.
Core (CUDA)
CUDA Core: a general-purpose arithmetic unit for scalar and element-wise operations.
Not every operator is handled by Tensor Cores.
Core (Tensor)
Tensor Core: a specialized unit optimized for mixed-precision matrix multiply-accumulate (MMA).
GEMM and most of the main inference compute rely heavily on it.
Streaming Multiprocessor (SM)
SM: a GPU compute unit that contains cores and caches.
Kernels are scheduled onto SMs, using large numbers of threads to achieve parallelism.
Thread
Thread: the smallest unit of execution on a GPU.
A kernel usually launches a huge number of threads to cover data-parallel work.
Special Function Unit (SFU)
SFU: a specialized unit that accelerates specific math operations such as sine and cosine.
It takes special functions off the CUDA Cores.
L0/L1/L2 caches (GPU)
The GPU's on-chip cache hierarchy, used for instructions, shared memory, global cache, and more.
Whether data hits a closer cache affects a kernel's real-world performance.
CUDA
NVIDIA's GPU programming model and platform, covering kernels, graphs, memory, and execution.
Inference software, compilers, and libraries usually reach the GPU through CUDA.
CUDA driver
CUDA driver: the low-level interface between applications and GPU hardware that manages memory and execution.
Handles system-level resources and device interaction.
CUDA runtime
CUDA runtime: the developer-facing API for launching kernels and managing memory.
Closer to application code and runtime calls than the driver is.
CUDA kernel
CUDA kernel: a user-defined function that runs in parallel on the GPU.
Operator performance optimization often comes down to kernel design in the end.
CUDA graph
CUDA graph: organizes kernels and other GPU operations into a DAG to optimize repeated workflows.
Cuts repeated launch overhead; well suited to inference with fairly stable shapes and flow.
Basic Linear Algebra Subprograms (BLAS)
BLAS: a standard interface for basic linear algebra operations.
Core operators like GEMM are often called through BLAS or its GPU implementations.
cuBLAS
CUDA's BLAS implementation, offering high-quality basic primitives such as GEMM.
It's a key library for optimizing matrix multiplication on NVIDIA GPUs.
cuDNN
CUDA's library of deep neural network primitives.
Provides high-performance implementations of common neural network operations.
General matrix-matrix multiplication (GEMM)
GEMM: general matrix-matrix multiplication, part of BLAS and a key operation in inference.
A huge number of linear layers ultimately boil down to matrix multiplication.
Matmul
Short for matrix multiplication.
Common in speech or code; GEMM is the more specific term that carries matrix shapes and implementation context.
CuTe
A C++ template library in the NVIDIA ecosystem that abstracts tiled tensor operations, helping compose precision-aware optimized GEMMs and fused kernels.
Closer to a developer tool for writing low-level, high-performance kernels.
CUTLASS
A CUDA C++ template library that provides building blocks for high-performance, architecture-tuned GEMM and related kernels.
Commonly used to assemble or tune matrix computations yourself.
DeepGEMM
An efficient GEMM kernel library created by the DeepSeek team, with strong performance on FP8.
Represents specialized kernel optimization for a specific precision and hardware.
FlashAttention
A series of optimized attention kernels that reduce memory traffic.
Examples from the book: FlashAttention 3 targets Hopper, and FlashAttention 4 targets Blackwell.
Kernel fusion
Kernel fusion: combining multiple kernels into one to reduce memory round-trips for intermediate results.
Often trades a more complex kernel for fewer reads, writes, and launch overheads.
FLOPS
Floating-point operations per second, usually measured in a Tensor Core context.
It's peak compute capability, not the same as real model throughput.
Arithmetic intensity
Arithmetic intensity: how many operations get done per byte moved.
Compared against the hardware's ops:byte, it tells you whether a kernel leans compute-bound or memory-bound.
Ops:byte ratio (GPU)
The peak number of operations per byte of memory bandwidth a GPU delivers at a given precision.
Used together with an algorithm's arithmetic intensity to help locate bottlenecks.
Roofline model
Roofline model: plots arithmetic intensity against memory bandwidth and compute ceilings on a single chart.
Used to decide whether optimization should focus on compute or memory bandwidth first.
Compute-bound
Compute-bound: performance is limited mainly by available FLOPS rather than memory bandwidth.
Prefill and image/video generation often sit closer to this side.
PyTorch compile (torch.compile)
A mechanism for GPU-specific graph capture, kernel selection, and fusion.
Caching compiled results can reduce cold start.
PyTorch Profiler
A developer tool that measures CPU/GPU time and memory per operation.
Diagnose bottlenecks from profile evidence, not just intuition when tuning.
5. Precision, quantization, and model compression
Term
What it means
What to look at in an inference system
Floating-point data formats
Floating-point formats: FP16, FP8, FP4, and so on, using an exponent–mantissa structure with high dynamic range.
Look at precision, VRAM, bandwidth, Tensor Core support, and quality loss together.
BF16
A 16-bit floating-point format with a wider exponent and greater dynamic range than FP16, so it handles outliers better.
Common in training, and used for some inference too.
Integer data formats
Integer formats: INT8, INT4, and the like, with a fairly limited dynamic range.
They usually save storage and bandwidth, but you have to watch out for quantization error.
Dynamic range (quantization)
Dynamic range: the span of absolute values a given number format can represent.
Floating-point formats generally offer greater dynamic range than integer formats of the same byte size.
Microscaling formats
Microscaling formats: MXFP8, MXFP4, NVFP4, and others, which apply scale factors to small blocks.
Blockwise scaling preserves as much precision as possible at low bit widths.
NVFP4
NVIDIA's 4-bit floating-point microscaling format, which uses two scale factors and block-level quantization with a block size of 16.
A concrete example of where low-precision hardware and formats are heading.
Quantization (post-training)
Post-training quantization: lowering the precision of weights, activations, and possibly the KV cache.
The goal is to cut compute and bandwidth, but you need quality and performance baselines to verify the result.
Quantization-aware training
Quantization-aware training: computing quantization scales and optimizing weights together so the final model is suited to low-precision deployment.
It generally adapts to quantization error more actively than plain post-training quantization, but it costs you training time.
Weights-only quantization
Weights-only quantization: lowering precision only for the weights of linear layers, while the KV cache, attention, and the rest stay at higher precision.
Quality usually holds up more reliably, but the performance gains may be more limited.
Scale factor (quantization)
Quantization scale factor: the multiplier that maps low-precision values back to the original numeric range.
How finely the scale is computed, and how, affects both the error and the kernel implementation.
Sparsity (FLOPS)
Sparsity: for example, 2:4 structured sparsity, where half the values are zero and Tensor Cores can skip the zero multiplications.
The book cautions that most inference is still dense, so don't assume sparsity is kicking in.
Fine-tuning
Fine-tuning: adapting a pretrained foundation model to a specific domain.
It can let a smaller model hit your target quality, which lowers deployment costs.
LoRA
Low-rank adaptation: a lightweight fine-tuning method that makes only small changes to a model.
Swapping among thousands of LoRAs on a single base model brings caching and routing headaches.
Distillation
Distillation: having a smaller student model mimic the probability distribution of a larger teacher, not just its final output.
It keeps some of the teacher's behavior with fewer parameters, making it one route to compression and lower cost.
6. Parallelism, Interconnects, and Distributed Inference
Term
What it means
What to watch in an inference system
Model parallelism (overview)
Model parallelism, broadly: splitting the work across multiple GPUs via TP, EP, PP, and so on.
Which one you pick depends on model size, topology, latency targets, and throughput targets.
Tensor Parallelism (TP)
Tensor parallelism: splitting tensor operations across multiple GPUs on the same node.
Latency for a single user is usually good, but it requires frequent all-reduce synchronization.
Pipeline Parallelism (PP)
Pipeline parallelism: slicing the model's layers into stages spread across GPUs.
It works for dense models across multiple nodes, but pipeline bubbles leave some GPUs waiting.
Expert Parallelism (EP)
Expert parallelism: sharding a MoE's experts across multiple GPUs, with each GPU holding several complete experts.
It raises overall throughput with relatively low inter-GPU communication overhead.
Mixture of Experts (MoE)
Mixture of experts: splitting the weights of linear layers into multiple sparse experts, with a router activating only some of them on each pass.
Total parameter count can be huge, but any single computation triggers only a fraction of the experts.
Multi-node inference
Multi-node inference: scaling out to two or more nodes when a single 8-GPU node doesn't have enough VRAM.
You need the right parallelism strategy; too much all-to-all over InfiniBand will slow TP down.
InfiniBand
The interconnect between nodes, used to scale training and inference across multiple nodes.
Bandwidth is usually higher than Ethernet, but well below intra-node NVLink.
NVLink
A high-speed point-to-point communication layer between GPUs.
The book cites examples of up to 1800 GB/s on Blackwell and up to 900 GB/s on Hopper; the exact figure depends on the product and link configuration.
NVSwitch
An all-to-all communication layer built on top of NVLink that coordinates all GPUs within a node.
It changes the communication topology and scalability inside a multi-GPU node.
PCIe (GPU form factor)
PCIe GPU form factor: connected through a standard PCI Express slot.
Base specs and interconnect options are usually fewer than on the SXM version of the same model.
SXM (GPU form factor)
SXM GPU form factor: a socketed module that supports higher-bandwidth connections and stronger power delivery.
Common in inference clusters for high-spec GPUs with strong interconnects.
Multi-Instance GPU (MIG)
Multi-Instance GPU: partitioning a larger GPU into as many as eight memory slices and seven compute slices.
Good for isolating small workloads, at the cost of some flexibility with the whole card.
Control plane (multi-cloud)
Multi-cloud control plane: the global orchestrator that deploys models and allocates resources.
It makes decisions, schedules, and manages capacity, but doesn't necessarily handle actual requests.
Workload plane (multi-cloud)
Multi-cloud workload plane: the separate clusters that actually run inference and handle requests.
The control plane sends work here to be executed.
Multi-cloud capacity management
Multi-cloud capacity management: global scheduling that places workloads across multiple cloud providers and regions.
The aim is to treat a heterogeneous pool of GPUs as schedulable resources while handling capacity and regional constraints.
Bin packing (multi-cloud)
Multi-cloud bin packing: treating GPU pools across different clouds, regions, and clusters as one resource pool to fill with tasks.
It requires capacity management infrastructure that hides the hardware heterogeneity.
Data sovereignty
Data sovereignty: legal constraints on where model inputs and outputs are processed and stored.
It can directly determine region selection, routing, and how you deploy across clouds.
Hyperscaler
Hyperscalers such as AWS and GCP.
Huge resource scale and general-purpose products; the counterpart to GPU-focused neoclouds.
Neocloud
Specialized GPU-focused cloud providers such as CoreWeave and Nebius.
Often used as a supplementary source of multi-cloud GPU capacity.
7. Production Deployment, Reliability, and Service Governance
Term
What it means
What to watch in an inference system
Active-active
Active-active: multiple regions or clusters serve live traffic at the same time, so if one plane fails, the others keep serving.
High availability, but it requires coordinating routing, state, and capacity across locations.
Active-passive
Active-passive: a hot standby cluster or region stays ready but idle, taking over if the primary fails.
Structurally simpler, but you have to weigh standby utilization against failover speed.
Autoscaling
Autoscaling: automatically adding or removing model replicas based on traffic or utilization.
The goal is to match capacity to demand while holding latency SLAs and cutting waste.
Autoscaling window
Autoscaling window: the rolling time range that triggers a scale-up or scale-down.
A longer window is more stable; a shorter one reacts faster to sudden spikes.
Cold start
Cold start: the time from scaling a replica up from zero to its first successful response, including GPU provisioning, container startup, model loading, and engine compilation.
It's the main cost of scale-to-zero and directly affects queueing and first-request latency.
Scale to zero
Scale to zero: shut down all replicas when idle and spin them up on demand when requests arrive.
Saves cost, but depends on fast cold starts and reliable queueing.
Blue-green deployment
Blue-green deployment: maintain two parallel production environments and shift traffic between blue and green.
Enables zero-downtime releases and fast rollbacks.
Canary deployment
Canary deployment: send a small share of live traffic to the new version first, watch its stability and performance, then ramp up gradually.
A way to validate progress on real traffic.
Shadow traffic
Shadow traffic: mirror real production requests to a candidate deployment, but don't use the candidate's results to serve users.
Lets you evaluate a new version without affecting the main path.
Routing (inference)
Inference routing: send requests to the right replica based on load, KV cache, available LoRAs, and sequence shape.
A good router does more than round-robin — it also uses the model's runtime state.
gRPC
gRPC: a structured, schema-first, bidirectional streaming protocol.
Good for well-defined service-to-service communication; schema validation adds some overhead.
WebSocket
WebSocket: a lightweight, bidirectional streaming protocol.
Well suited to unstructured audio chunks and real-time user experience.
Service Level Agreement (SLA)
SLA: a contractual commitment the system makes about metrics such as latency, throughput, and availability.
An outward-facing promise, usually backed by observability data.
Service Level Objective (SLO)
SLO: an internal target the system sets to meet or exceed an SLA.
Internal targets should generally leave headroom above the external SLA.
Docker
Docker: packages an inference service and its dependencies into a standardized runtime unit using containers.
Makes deployment environments more reproducible, but doesn't automatically solve GPU driver or performance issues.
Dockerfile
Dockerfile: a human-readable, machine-executable instruction file for building a container image.
Pins down runtime dependencies, startup commands, and the environment.
8. Inference Engines, Runtimes, and Engineering Tools
Term
What it means
What to look at in an inference system
vLLM
A widely adopted inference engine with broad model and hardware support and mature defaults.
Common capabilities include continuous batching, KV management, quantization, and serving APIs.
SGLang
A fast inference engine with a flexible frontend/backend and strengthened MoE support.
Suited to scenarios that need structured workflows and high-performance serving.
TensorRT
NVIDIA's runtime optimized for high-performance inference.
Combines optimizations such as fused kernels and quantization.
TensorRT-LLM
An LLM inference engine built by NVIDIA, offering a Python API and supporting both TensorRT engines and a PyTorch backend.
Integrates fused kernels, quantization, and speculative decoding.
It handles model serving orchestration, protocols, and backend integration — it isn't a single kernel library.
ONNX
An intermediate representation for models and its runtime ecosystem.
Used to exchange model representations across different frameworks or runtimes.
NIM
Pre-packaged, containerized microservices that NVIDIA provides for specific models.
Lowers the barrier to packaging model serving, but you still need to watch the underlying hardware and runtime behavior.
NVIDIA Dynamo
NVIDIA's open-source distributed serving platform, aimed at KV reuse, disaggregated inference, and multi-GPU/multi-node orchestration.
It sits at the serving orchestration layer, connecting model execution and cluster resources.
Transformers (library)
A reference implementation library for LLMs and other Transformer models.
Commonly used for model loading, configuration, inference logic, and research prototypes.
Diffusers (library)
A reference implementation library for image and video generation pipelines.
Strings components like text encoding, denoising, and the VAE into a generation pipeline.
ComfyUI
A workflow tool for assembling image generation pipelines.
Uses nodes to combine components such as the base model, refiner, LoRA, and ControlNet.
9. Diffusion, Multimodal, Vision, and Speech Inference
Term
What it means
What to look at in an inference system
Image generation pipeline
Image generation pipeline: typically composed of several models, such as a text encoder, an iterative denoiser, and a VAE.
It's not as simple as one forward pass of one model — multiple stages run in sequence.
Denoising model
Denoising model: the core of a diffusion pipeline, repeatedly refining latent noise into an image or video.
Inference steps, latent size, and kernels directly affect latency.
Iterative denoising (diffusion)
Iterative denoising: starting from noise, generating an image or video step by step in latent space.
Each step runs denoising computation, and the number of steps is a major performance lever.
Latent space (images/videos)
A low-dimensional latent space for images/videos where denoising usually happens, for example a 128x128 representation.
Computing in latent space uses fewer resources than computing directly on pixels.
VAE (variational autoencoder)
Variational autoencoder: decodes latents into pixel space; during training it can also encode pixels back into latents.
It's the key boundary in a diffusion pipeline between the latent and the final image.
CLIP (text encoder)
The CLIP text/image encoder, commonly used in early image pipelines such as SDXL.
Modern systems often replace or augment prompt understanding with a stronger, full LLM.
Classifier-free guidance
Classifier-free guidance: at each step, balancing unconditional and prompt-conditioned denoising.
Lower guidance is more creative; higher guidance follows the prompt more closely, but can affect quality and speed.
Few-step image generation
Few-step image generation: producing a usable image in eight steps or fewer.
Can be 80–90% faster, but usually comes with a clear quality trade-off — good for real-time scenarios.
Latent consistency
Latent consistency: a few-step strategy that directly predicts the target latent and can be refined repeatedly.
Very fast, but fidelity is usually lower than full diffusion.
SDXL
A representative early diffusion image pipeline: base + refiner + CLIP.
Modern systems usually keep the multi-stage structure but swap in stronger components.
Omni-modal
Omni-modal: accepts many kinds of input — text, images, video, audio — and produces many kinds of output too.
The key point is that neither input nor output modalities are limited to text.
Vision-language model (VLM)
Vision-language model: takes images/video and a text prompt as input and outputs text.
Inference has to account for visual encoding, cross-modal attention, and text decoding.
Encoder
Encoder: a network that turns raw input into an internal representation, such as Whisper's audio feature encoder.
Paired with a decoder in encoder-decoder models, it can serve as an early stage in a multimodal pipeline.
Automatic Speech Recognition (ASR)
Automatic Speech Recognition: a transcription model that takes audio input and produces text output, such as Whisper.
Decoder work often dominates runtime and can benefit from LLM-style optimizations and in-flight batching.
Voice activity detection (VAD)
Voice activity detection: splits an audio stream or file into segments containing speech.
A lightweight model used in ASR preprocessing to cut wasted compute on segments with no speech.
Diarization
Diarization: answers 'who spoke when', splitting audio by speaker.
Often paired with VAD to add speaker boundaries to ASR output.
Neural audio codec
Neural audio codec: compresses audio into tokens, which a matching decoder turns back into audio.
Turns audio generation into a token-level modeling problem.
SNAC (audio decoder)
A high-performance audio decoder path, often used alongside TTS token streams.
Its job is to turn a stream of audio tokens back into playable audio.
10. Speculative Decoding and Sampling Acceleration
These techniques all share one goal: make a single expensive forward pass through the target model confirm as many tokens as possible.
Term
What it means
What to watch in an inference system
Speculative decoding
Speculative decoding: draft tokens are generated first, then verified in a batch by the target model so a single forward pass can accept several tokens at once.
The key metrics are draft cost, acceptance rate, and how many tokens end up accepted per step.
EAGLE (speculation)
EAGLE: a small, specially trained draft model that consumes hidden states and proposes multiple tokens.
The aim is to make the draft align more closely with the target, raising the acceptance rate.
Medusa (speculation)
Medusa: adds extra decoder heads through fine-tuning so each forward pass produces several draft tokens.
It doesn't necessarily need a separate draft model, but it does require extra training or changes to the model architecture.
Lookahead decoding
Lookahead decoding: builds n-grams during inference so draft tokens can be predicted without a separate model.
It exploits patterns already present in the generation to cut down on serial dependencies.
N-gram speculation
N-gram speculation: uses n-grams observed during prefill to propose longer draft sequences at decode time.
Especially effective for code completion, since code tends to repeat itself and follow stable local patterns.
11. Retrieval, Embeddings, and Model Quality Evaluation
Term
What it means
What to watch in an inference system
Embedding model
Embedding model: encodes text or images into fixed-dimension vectors for semantic similarity.
Commonly powers RAG, search, and agent memory; modern versions often use an LLM backbone.
Matryoshka representations (embeddings)
Matryoshka representations: nested vector representations where the front of the vector carries more semantic weight, so dimensions can be truncated as needed.
Offers a tunable trade-off between vector size, retrieval cost, and quality.
Vector database
Vector database: stores and queries the semantic vectors produced by embeddings.
The retrieval layer of RAG; performance hinges on indexing, recall, filtering, and storage.
Vector similarity
Vector similarity: measures how close two vectors are using formulas such as cosine similarity.
Similar vectors usually mean similar semantics, but similarity is not factual correctness.
Retrieval-augmented generation (RAG)
Retrieval-augmented generation: an application pattern that fetches extra context for an LLM beyond the prompt itself.
End-to-end latency must include retrieval, reranking, context assembly, and the extra prefill.
Benchmark (intelligence)
Intelligence benchmark: measures how well a model answers questions or takes appropriate actions, e.g. MMLU.
It gauges model capability, not actual product quality.
Benchmark (performance)
Performance benchmark: measures the latency and throughput of an inference service under a given model and workload.
Hardware, ISL/OSL, concurrency, batch size, and measurement methodology must all be fixed.
Evals
Evals: task-specific tests that simulate real use cases to measure model intelligence in production scenarios.
Closer to the business than generic benchmarks, but they demand more work to maintain the data and criteria yourself.
Elo (quality meta-metric)
Elo as a quality meta-metric: compares model quality by pairwise win rates.
A directional signal only; it shouldn't replace task-specific metrics or human review on its own.
Goodhart's Law
Goodhart's Law: once a measure becomes a target, it stops being a good measure.
When optimizing TPS, TTFT, or a benchmark, keep an eye on quality, cost, and the real user experience at the same time.
12. GPU Architectures, Chips, and System Models
This section collects hardware terms that come with lots of names but shouldn't be confused with algorithmic concepts. The figures and roadmaps here are a snapshot of the book's edition.
Term
What it means
What to watch in an inference system
Ada Lovelace (architecture)
NVIDIA's graphics-oriented GPU architecture, released around the same time as Hopper; suited to small models and cost-sensitive workloads.
The book considers it a poor fit for large-scale LLM inference.
Ampere (architecture)
An earlier NVIDIA GPU architecture, still used in legacy or small-scale deployments.
When comparing it with Hopper and Blackwell, focus on cost, bandwidth, and efficiency at scale.
Hopper (architecture)
NVIDIA's 2022 GPU architecture, supporting FP8 and asynchronous programming features.
One of the target architectures for FlashAttention 3.
Blackwell (architecture)
NVIDIA's GPU generation from late 2024, supporting FP4, MXFP8, MXFP4, NVFP4, and high memory bandwidth.
In an inference context, the focus is on low-precision formats, bandwidth, and serving large models.
Rubin (architecture)
NVIDIA's next-generation architecture as described in the book, introducing HBM4 and CPX for compute-bound workloads.
A future/roadmap term as of the book's publication; check the latest official specs before relying on it.
Feynman (architecture)
A future NVIDIA architecture following Rubin, described in the book with limited detail.
Treat it as a roadmap name only; don't infer shipping capabilities from it.
B200
A Blackwell data center GPU; the book describes it as having 192 GB of VRAM, 8 TB/s of bandwidth, and 5 petaFLOPS of FP8.
These are the book's snapshot figures; actual SKUs and measurement methodology need to be verified separately.
B300
A Blackwell data center GPU; the book describes it as having 288 GB of VRAM, 8 TB/s of bandwidth, and 5 petaFLOPS of FP8.
Mainly useful for understanding how VRAM, bandwidth, and precision affect inference.
Grace CPU
NVIDIA's ARM CPU, working alongside GPUs over a high-bandwidth chip-to-chip interconnect.
Suited to systems that need fast CPU/GPU exchange, such as KV offload.
Vera CPU
An NVIDIA ARM CPU described in the book as succeeding Grace in the Rubin GPU generation.
This is a roadmap architecture name; specific product capabilities need to be verified per release.
GB200
An NVIDIA superchip combining a Grace CPU with a B200 GPU, linked chip-to-chip over high-bandwidth NVLink.
Suited to techniques that benefit from NVLink-C2C, such as KV cache offload and LoRA swapping.
GH200
An NVIDIA superchip combining a Grace CPU with an H200 GPU, linked chip-to-chip over high-bandwidth NVLink.
Similar to GB200, with the emphasis on high-bandwidth CPU/GPU cooperation.
NVL72
A rack-scale Blackwell system described in the book as interconnecting 72 GPUs and 36 CPUs.
Built for very large models and extremely high-throughput serving; system-level topology matters more than single-card specs.
This article compiles the terminology from Appendix A of Inference Engineering by Philip Kiely, published by Baseten Books. The original book orders terms alphabetically; this article regroups them along the inference system stack. For hardware models, software versions, and future architecture roadmaps, defer to the book's edition and official sources.
Inference Engineering: A Glossary of Inference Engineering Terms, Organized by System Pipeline
Once a generative AI app actually ships, the questions stop being just "can the model answer?" and start including how the model gets loaded, how requests are queued, how the GPU computes and moves data, and how the service stays stable under real traffic.
This article is compiled from
Inference Engineering by Philip Kiely. The book is published by
Baseten Books, with an official
online reading version. The terms below come mainly from the book's
Appendix A: Inference Glossary. The book lists them alphabetically; here I've regrouped them by inference system pipeline so you can cross-reference them while writing code, reading source, or looking at performance metrics.
The book spans everything from CUDA, GPUs, and inference engines to KV cache, quantization, parallelism, diffusion models, speech, multi-cloud deployment, and production serving. I got through it in two days — and Appendix B is a great indexing resource too.
1. Models, Applications, and Data Representation
This group answers "what a model is, how an application calls it, and how text becomes a representation the model can process."
Term
What it means
What to look at in an inference system
Activation function
Activation function: a mostly differentiable nonlinear function inserted between linear layers, e.g. ReLU.
Without nonlinearity, a multilayer network collapses into a single matrix multiplication.
Generative AI
Generative AI: learns patterns from data and generates new text, images, audio, video, or code.
The emphasis is on "generating new content," as opposed to traditional ML, which only classifies or predicts.
Machine learning (ML)
Machine learning: learns predictive models from data, e.g. classification and trend forecasting.
In the glossary, it serves as the contrast to generative AI.
Foundation model
Foundation model: trained on broad data and usable as the base for many downstream tasks.
You can prompt it directly, or fine-tune it into a domain-specific model.
Open model
Open model: model weights are freely available, e.g. Llama, DeepSeek, Whisper.
Visible weights usually mean more control over deployment, quantization, and hardware choices.
Closed model
Closed model: proprietary model whose weights are not available, e.g. GPT, Claude, Gemini.
Typically used through an API — what you control is requests, routing, and cost, not the underlying weights.
Transformer
Transformer: the foundational network architecture behind generative AI.
Attention, Q/K/V, KV cache, and most inference optimizations revolve around it.
Causal language model (CLM)
Causal language model: a decoder-only Transformer that predicts the next token using only the preceding context.
Autoregressive LLM inference is essentially running a CLM's next-token prediction loop.
Large Language Model (LLM)
Large language model: takes a text prompt and generates a new sequence of text.
Typical families include GPT, Claude, Llama, and DeepSeek.
LLM
Short for `Large Language Model`.
Same concept as the row above; in engineering docs it's usually written simply as LLM.
Generative Pretrained Transformer (GPT)
Generative Pretrained Transformer: the family of text-generation LLMs created by OpenAI.
It's the name of a model family, not a catch-all for all LLMs.
Agent
Agent: an AI application that doesn't just answer questions but also calls tools and takes action.
A single user request may trigger multiple inference calls, multiple models, and multiple modalities.
AI-native application
AI-native application: a product whose core experience and value depend on generative models.
Modality, latency budget, unit economics, and usage patterns all feed back into the inference architecture.
Application Programming Interface (API)
API: a structured interface for sending requests and receiving responses.
Inference engines typically expose model querying through an API.
Inference
Inference: serving an AI model in production.
The emphasis here is on serving, not just "running a single forward pass."
Inference engine
Inference engine: a high-performance runtime supporting optimizations such as batching, caching, quantization, and speculation.
vLLM, SGLang, and TensorRT-LLM all fall into this category.
Prompt
Prompt: the instruction given to the model; for diffusion models it may also include a negative prompt, step count, and guidance parameters.
Prompt length directly affects prefill, TTFT, and KV cache usage.
Chat template
Chat template: serializes roles, delimiters, and sequence start/end tokens into input as the model requires.
Change the template for the same conversation and the actual token sequence may differ.
Token
Token: the basic unit of text an LLM processes — essentially an integer representing a string fragment.
Latency, throughput, context window, and billing are usually measured in tokens.
Tokenizer
Tokenizer: performs deterministic conversion between strings and token sequences.
Different models use different tokenizers; a more efficient tokenizer can lower end-to-end latency.
Vocabulary
Vocabulary: the full set of tokens a model uses to represent data.
Vocabulary size affects the logits vector and tokenization behavior.
Input sequence
Input sequence: the tokens in a request handed to the model, processed during prefill.
The longer the input, the higher the cost of prefill computation and KV cache setup.
Input Sequence Length (ISL)
Input sequence length: the number of input tokens in a single request.
Together with OSL, it describes the length shape of a workload.
Output sequence
Output sequence: the tokens the model generates during decode.
Output length determines how long decode runs and what the user sees streamed.
Output Sequence Length (OSL)
Output sequence length: the number of output tokens generated by a single request.
A long OSL tends to amplify decode, KV cache, and queue pressure.
Context window
Context window: the total cap on input, reasoning, and output tokens a model can process in a single request.
It's both the ceiling on model capability and a constraint on VRAM, KV cache, and scheduling.
Function calling
Function calling, also known as tool calling / tool use: the model picks a function from a given set and returns structured arguments.
The server must validate the schema, execute the tool, and feed the result back to the model.
Structured output
Structured output: model output that follows a specified schema.
The book stresses achieving this with generation constraints such as logit bias, not just by agreeing on a prompt.
Logit biasing
Logit biasing: adjusting or constraining token probabilities before sampling to steer output such as JSON or tool calls.
It sits after logits are produced and before the final sampling step.
2. Request Lifecycle, Decoding, and Performance Metrics
These terms cover "what happens after a request comes in, and how should speed be measured."
Term
What it means
What to look at in an inference system
Autoregressive token generation
Autoregressive token generation: each new token depends on the tokens already generated.
This naturally forms a step-by-step decode loop and is the backdrop for KV cache and speculative decoding.
Pretraining
Pretraining: large-scale training on broad corpora to produce a foundation model.
It happens before serving; inference engineering typically consumes the weights it produces.
Training
Training: learning model weights from data through backpropagation and optimization.
Training is compute-heavy and usually relies on large-scale GPU clusters.
Prefill
Prefill: the LLM inference phase that processes the input sequence in one pass and builds the KV cache.
It is usually compute-bound, and a long ISL can significantly increase TTFT.
Decode
Decode: the phase in which the autoregressive loop generates one token at a time.
This phase tends to be constrained by memory bandwidth and KV reads; each forward pass usually adds just one token.
Sampling (decode)
Sampling: selecting the next output token from the generated logits.
Common strategies include greedy/argmax, temperature, top-k, and top-p.
Logits
Logits: the vector of unnormalized scores the model produces for every token in the vocabulary on each forward pass.
Sampling, temperature, and logit bias all operate on it or on a transformation of it.
Softmax
Softmax: turns attention scores into probabilities, and also normalizes the logits during decode into a probability distribution.
It bridges the two stages of scoring and probability-based sampling.
Temperature
Temperature: controls how random token selection is; low values are more deterministic, high values more diverse.
It changes the sampling distribution; it does not change the model's underlying capabilities.
Batch
Batching: processing multiple inputs at once.
It improves device utilization, but differences in request length lead to waiting or padding.
Batch sizing
Batch size: a core lever for trading off latency against throughput.
A larger batch usually raises total throughput, but it worsens latency for each individual user.
Dynamic batching
Dynamic batching: fires once the batch is full or a short timer expires.
It balances utilization and latency stability; for LLMs it has largely been replaced by continuous batching.
Continuous batching (in-flight)
Continuous batching / in-flight batching: interleaves requests at the token level so GPU slots always have work to do.
This matters especially for online LLM requests that vary in length and arrive continuously.
In-flight batching
Another name for `continuous batching`.
The key point is that requests can join, finish, and exit while a batch is still executing.
Chunked prefill
Chunked prefill: splitting a long input into chunks and overlapping them with decode or other work.
It keeps one very long sequence from monopolizing resources for an extended period.
Queue (request)
Request queue: holds traffic that exceeds current capacity until autoscaling spins up new replicas.
Queue length and wait time are signals of insufficient capacity, cold starts, or traffic spikes.
Online inference
Online inference: serving requests in real time, prioritizing strict latency budgets.
Watch TTFT, ITL, P99, and availability.
Offline inference
Offline inference: processing large jobs asynchronously in batches, prioritizing throughput and cost.
You can trade per-request latency for higher device utilization.
Local (edge) inference
Local / edge inference: running inference on end-user devices such as phones and laptops.
It cuts out network round trips, but is limited by on-device compute, memory, and power.
Latency percentiles
Latency percentiles: using P50, P90, P95, P99, and so on to see the distribution behind average-to-tail latency.
P99 is often closer to the worst user experience than the average is.
Time to first byte (TTFB)
Time to first byte: the time from request start to receiving the first output byte.
It is shaped by the network, server-side processing, and the streaming protocol together.
Time to first token (TTFT)
Time to first token: the time from request start to receiving the first output token.
This is critical for interactive scenarios like chat and code completion, and it usually includes prefill.
Inter-token latency (ITL)
Inter-token latency: the interval between consecutive generated tokens during decode.
The smaller the ITL, the smoother the streamed output feels to users; 2 ms, for example, corresponds to roughly 500 TPS.
Perceived TPS
Perceived TPS: the number of tokens per second a single user actually sees in the streamed output.
It reflects the individual user experience more closely than aggregate throughput does.
Tokens per second (TPS)
Tokens per second: the number of tokens streamed to the end user each second; in the glossary this refers to perceived TPS.
Be clear about whether you mean single-user TPS or aggregate throughput.
Throughput
Throughput: the total work completed per unit of time, such as total tokens/s.
Higher aggregate throughput does not mean better TPS or TTFT for each user.
Baselines
Baselines: carefully recorded performance and quality measurements taken before optimization.
Without a baseline, you cannot reliably attribute gains or regressions to an optimization.
Load testing
Load testing: sending sustained high traffic to probe throughput limits, queue behavior, and scaling.
Pair it with realistic ISL/OSL distributions rather than uniform synthetic requests alone.
Jitter traffic (bench)
Jitter traffic: adding randomness to request arrival times and sequence shapes.
This comes closer to real workloads than tidy, fixed-interval traffic.
Real-time factor (RTF)
Real-time factor: a speed metric for ASR transcription; if one hour of audio finishes transcribing in six seconds, the RTF is 600.
It is well suited to expressing audio processing speed relative to the original playback duration.
3. Attention, the KV Cache, and Context Optimization
Term
What it means
What to look at in an inference system
Attention
Attention: the Transformer mechanism that links the current token to historical tokens through Q/K/V projections and softmax.
It is heavy in both compute and memory traffic, making it a prime target for inference optimization.
Head (attention)
Attention head: a single independent attention computation within a layer.
The multi-head structure lets the model capture relationships from different subspaces.
Cross-attention
Cross-attention: the Q of one sequence is conditioned on the K/V of another sequence.
It shows up often in text-conditioned image generation, multimodal models, and diffusion denoising pipelines.
KV cache
KV cache: stores the K/V tensors for each historical token so history projections are not recomputed at every step.
It trades VRAM for compute; the longer the context, the larger the cache and the more there is to read at each step.
Prefix caching
Prefix caching: reusing the KV of shared prefixes across requests to skip redundant prefill.
It can significantly improve TTFT for code completion, multi-turn conversations, and agent systems.
Cache-aware routing
Cache-aware routing: sending requests to replicas that already hold the matching prefix or the required LoRA.
Routing considers cache hits and local state, not just load.
PagedAttention
PagedAttention: storing KV blocks in fixed-size pages to manage long-context memory more effectively.
Its typical benefits are less fragmentation and higher KV utilization; do not conflate it with ordinary OS paging.
Rotary positional embeddings (RoPE)
Rotary positional embeddings: representing position with learnable rotations, which helps with long-context extrapolation.
There is a trade-off between long-context capability and attention memory requirements.
Context Parallelism (CP)
Context parallelism: replicating weights across GPUs while splitting the attention context.
It suits models with extremely large contexts or video latent spaces.
Ring attention
Ring attention: a CP mechanism in which GPUs pass partial attention results around in a ring.
It enables longer contexts by reducing all-to-all pressure.
Disaggregation
Disaggregated inference: splitting prefill and decode into engines that scale independently and use different hardware resources.
It lets two very different bottlenecks scale separately, but adds transfer and scheduling complexity.
4. GPUs, the Memory Hierarchy, and the CUDA Software Stack
This group of terms explains where the model actually computes, how data moves, and how kernels execute.
Term
What it means
What to look at in an inference system
Central Processing Unit (CPU)
CPU: a general-purpose processor suited to sequential workloads.
Mostly handles orchestration, scheduling, networking, and preprocessing; it usually doesn't do the main generation compute itself.
Graphics Processing Unit (GPU)
GPU: a highly parallel processor, now widely used for both training and inference in generative AI.
Look at memory capacity, bandwidth, compute units, and interconnect—not just the model name.
GPU node
GPU node: typically a standard chassis holding 8 GPUs connected via NVLink/NVSwitch.
It's the common building block for multi-GPU inference and intra-node parallelism.
Node
Node: a physical 8-GPU base unit with NVLink/NVSwitch; multiple nodes are then linked with InfiniBand.
Clarifies the boundary between intra-node and inter-node communication.
Instance (cloud)
Cloud instance: a configured virtual machine that includes GPU, CPU, RAM, storage, networking, and interconnect.
On the cloud, billing, deployment, and autoscaling usually happen at the instance level.
High-Bandwidth Memory (HBM)
High-bandwidth memory: the high-bandwidth memory used as VRAM in data center GPUs, such as HBM3, HBM3e, and HBM4.
It determines how much room weights, activations, and KV have, and how fast data can be fed in.
VRAM (device memory)
VRAM: the device memory on the GPU that holds weights, KV, and activations.
Capacity limits model size and KV headroom; bandwidth affects decode TPS.
Out-of-memory error (OOM)
Out-of-memory error: the GPU doesn't have enough VRAM to load weights or run inference.
Troubleshoot model precision, context length, batch, KV cache, and parallelism together.
Bandwidth
Bandwidth: how much data a memory or interconnect can move per second.
Decode is often bandwidth-bound; NVLink, HBM, and InfiniBand are bandwidth at different levels.
Core (CUDA)
CUDA Core: a general-purpose arithmetic unit for scalar and element-wise operations.
Not every operator is handled by Tensor Cores.
Core (Tensor)
Tensor Core: a specialized unit optimized for mixed-precision matrix multiply-accumulate (MMA).
GEMM and most of the main inference compute rely heavily on it.
Streaming Multiprocessor (SM)
SM: a GPU compute unit that contains cores and caches.
Kernels are scheduled onto SMs, using large numbers of threads to achieve parallelism.
Thread
Thread: the smallest unit of execution on a GPU.
A kernel usually launches a huge number of threads to cover data-parallel work.
Special Function Unit (SFU)
SFU: a specialized unit that accelerates specific math operations such as sine and cosine.
It takes special functions off the CUDA Cores.
L0/L1/L2 caches (GPU)
The GPU's on-chip cache hierarchy, used for instructions, shared memory, global cache, and more.
Whether data hits a closer cache affects a kernel's real-world performance.
CUDA
NVIDIA's GPU programming model and platform, covering kernels, graphs, memory, and execution.
Inference software, compilers, and libraries usually reach the GPU through CUDA.
CUDA driver
CUDA driver: the low-level interface between applications and GPU hardware that manages memory and execution.
Handles system-level resources and device interaction.
CUDA runtime
CUDA runtime: the developer-facing API for launching kernels and managing memory.
Closer to application code and runtime calls than the driver is.
CUDA kernel
CUDA kernel: a user-defined function that runs in parallel on the GPU.
Operator performance optimization often comes down to kernel design in the end.
CUDA graph
CUDA graph: organizes kernels and other GPU operations into a DAG to optimize repeated workflows.
Cuts repeated launch overhead; well suited to inference with fairly stable shapes and flow.
Basic Linear Algebra Subprograms (BLAS)
BLAS: a standard interface for basic linear algebra operations.
Core operators like GEMM are often called through BLAS or its GPU implementations.
cuBLAS
CUDA's BLAS implementation, offering high-quality basic primitives such as GEMM.
It's a key library for optimizing matrix multiplication on NVIDIA GPUs.
cuDNN
CUDA's library of deep neural network primitives.
Provides high-performance implementations of common neural network operations.
General matrix-matrix multiplication (GEMM)
GEMM: general matrix-matrix multiplication, part of BLAS and a key operation in inference.
A huge number of linear layers ultimately boil down to matrix multiplication.
Matmul
Short for matrix multiplication.
Common in speech or code; GEMM is the more specific term that carries matrix shapes and implementation context.
CuTe
A C++ template library in the NVIDIA ecosystem that abstracts tiled tensor operations, helping compose precision-aware optimized GEMMs and fused kernels.
Closer to a developer tool for writing low-level, high-performance kernels.
CUTLASS
A CUDA C++ template library that provides building blocks for high-performance, architecture-tuned GEMM and related kernels.
Commonly used to assemble or tune matrix computations yourself.
DeepGEMM
An efficient GEMM kernel library created by the DeepSeek team, with strong performance on FP8.
Represents specialized kernel optimization for a specific precision and hardware.
FlashAttention
A series of optimized attention kernels that reduce memory traffic.
Examples from the book: FlashAttention 3 targets Hopper, and FlashAttention 4 targets Blackwell.
Kernel fusion
Kernel fusion: combining multiple kernels into one to reduce memory round-trips for intermediate results.
Often trades a more complex kernel for fewer reads, writes, and launch overheads.
FLOPS
Floating-point operations per second, usually measured in a Tensor Core context.
It's peak compute capability, not the same as real model throughput.
Arithmetic intensity
Arithmetic intensity: how many operations get done per byte moved.
Compared against the hardware's ops:byte, it tells you whether a kernel leans compute-bound or memory-bound.
Ops:byte ratio (GPU)
The peak number of operations per byte of memory bandwidth a GPU delivers at a given precision.
Used together with an algorithm's arithmetic intensity to help locate bottlenecks.
Roofline model
Roofline model: plots arithmetic intensity against memory bandwidth and compute ceilings on a single chart.
Used to decide whether optimization should focus on compute or memory bandwidth first.
Compute-bound
Compute-bound: performance is limited mainly by available FLOPS rather than memory bandwidth.
Prefill and image/video generation often sit closer to this side.
PyTorch compile (torch.compile)
A mechanism for GPU-specific graph capture, kernel selection, and fusion.
Caching compiled results can reduce cold start.
PyTorch Profiler
A developer tool that measures CPU/GPU time and memory per operation.
Diagnose bottlenecks from profile evidence, not just intuition when tuning.
5. Precision, quantization, and model compression
Term
What it means
What to look at in an inference system
Floating-point data formats
Floating-point formats: FP16, FP8, FP4, and so on, using an exponent–mantissa structure with high dynamic range.
Look at precision, VRAM, bandwidth, Tensor Core support, and quality loss together.
BF16
A 16-bit floating-point format with a wider exponent and greater dynamic range than FP16, so it handles outliers better.
Common in training, and used for some inference too.
Integer data formats
Integer formats: INT8, INT4, and the like, with a fairly limited dynamic range.
They usually save storage and bandwidth, but you have to watch out for quantization error.
Dynamic range (quantization)
Dynamic range: the span of absolute values a given number format can represent.
Floating-point formats generally offer greater dynamic range than integer formats of the same byte size.
Microscaling formats
Microscaling formats: MXFP8, MXFP4, NVFP4, and others, which apply scale factors to small blocks.
Blockwise scaling preserves as much precision as possible at low bit widths.
NVFP4
NVIDIA's 4-bit floating-point microscaling format, which uses two scale factors and block-level quantization with a block size of 16.
A concrete example of where low-precision hardware and formats are heading.
Quantization (post-training)
Post-training quantization: lowering the precision of weights, activations, and possibly the KV cache.
The goal is to cut compute and bandwidth, but you need quality and performance baselines to verify the result.
Quantization-aware training
Quantization-aware training: computing quantization scales and optimizing weights together so the final model is suited to low-precision deployment.
It generally adapts to quantization error more actively than plain post-training quantization, but it costs you training time.
Weights-only quantization
Weights-only quantization: lowering precision only for the weights of linear layers, while the KV cache, attention, and the rest stay at higher precision.
Quality usually holds up more reliably, but the performance gains may be more limited.
Scale factor (quantization)
Quantization scale factor: the multiplier that maps low-precision values back to the original numeric range.
How finely the scale is computed, and how, affects both the error and the kernel implementation.
Sparsity (FLOPS)
Sparsity: for example, 2:4 structured sparsity, where half the values are zero and Tensor Cores can skip the zero multiplications.
The book cautions that most inference is still dense, so don't assume sparsity is kicking in.
Fine-tuning
Fine-tuning: adapting a pretrained foundation model to a specific domain.
It can let a smaller model hit your target quality, which lowers deployment costs.
LoRA
Low-rank adaptation: a lightweight fine-tuning method that makes only small changes to a model.
Swapping among thousands of LoRAs on a single base model brings caching and routing headaches.
Distillation
Distillation: having a smaller student model mimic the probability distribution of a larger teacher, not just its final output.
It keeps some of the teacher's behavior with fewer parameters, making it one route to compression and lower cost.
6. Parallelism, Interconnects, and Distributed Inference
Term
What it means
What to watch in an inference system
Model parallelism (overview)
Model parallelism, broadly: splitting the work across multiple GPUs via TP, EP, PP, and so on.
Which one you pick depends on model size, topology, latency targets, and throughput targets.
Tensor Parallelism (TP)
Tensor parallelism: splitting tensor operations across multiple GPUs on the same node.
Latency for a single user is usually good, but it requires frequent all-reduce synchronization.
Pipeline Parallelism (PP)
Pipeline parallelism: slicing the model's layers into stages spread across GPUs.
It works for dense models across multiple nodes, but pipeline bubbles leave some GPUs waiting.
Expert Parallelism (EP)
Expert parallelism: sharding a MoE's experts across multiple GPUs, with each GPU holding several complete experts.
It raises overall throughput with relatively low inter-GPU communication overhead.
Mixture of Experts (MoE)
Mixture of experts: splitting the weights of linear layers into multiple sparse experts, with a router activating only some of them on each pass.
Total parameter count can be huge, but any single computation triggers only a fraction of the experts.
Multi-node inference
Multi-node inference: scaling out to two or more nodes when a single 8-GPU node doesn't have enough VRAM.
You need the right parallelism strategy; too much all-to-all over InfiniBand will slow TP down.
InfiniBand
The interconnect between nodes, used to scale training and inference across multiple nodes.
Bandwidth is usually higher than Ethernet, but well below intra-node NVLink.
NVLink
A high-speed point-to-point communication layer between GPUs.
The book cites examples of up to 1800 GB/s on Blackwell and up to 900 GB/s on Hopper; the exact figure depends on the product and link configuration.
NVSwitch
An all-to-all communication layer built on top of NVLink that coordinates all GPUs within a node.
It changes the communication topology and scalability inside a multi-GPU node.
PCIe (GPU form factor)
PCIe GPU form factor: connected through a standard PCI Express slot.
Base specs and interconnect options are usually fewer than on the SXM version of the same model.
SXM (GPU form factor)
SXM GPU form factor: a socketed module that supports higher-bandwidth connections and stronger power delivery.
Common in inference clusters for high-spec GPUs with strong interconnects.
Multi-Instance GPU (MIG)
Multi-Instance GPU: partitioning a larger GPU into as many as eight memory slices and seven compute slices.
Good for isolating small workloads, at the cost of some flexibility with the whole card.
Control plane (multi-cloud)
Multi-cloud control plane: the global orchestrator that deploys models and allocates resources.
It makes decisions, schedules, and manages capacity, but doesn't necessarily handle actual requests.
Workload plane (multi-cloud)
Multi-cloud workload plane: the separate clusters that actually run inference and handle requests.
The control plane sends work here to be executed.
Multi-cloud capacity management
Multi-cloud capacity management: global scheduling that places workloads across multiple cloud providers and regions.
The aim is to treat a heterogeneous pool of GPUs as schedulable resources while handling capacity and regional constraints.
Bin packing (multi-cloud)
Multi-cloud bin packing: treating GPU pools across different clouds, regions, and clusters as one resource pool to fill with tasks.
It requires capacity management infrastructure that hides the hardware heterogeneity.
Data sovereignty
Data sovereignty: legal constraints on where model inputs and outputs are processed and stored.
It can directly determine region selection, routing, and how you deploy across clouds.
Hyperscaler
Hyperscalers such as AWS and GCP.
Huge resource scale and general-purpose products; the counterpart to GPU-focused neoclouds.
Neocloud
Specialized GPU-focused cloud providers such as CoreWeave and Nebius.
Often used as a supplementary source of multi-cloud GPU capacity.
7. Production Deployment, Reliability, and Service Governance
Term
What it means
What to watch in an inference system
Active-active
Active-active: multiple regions or clusters serve live traffic at the same time, so if one plane fails, the others keep serving.
High availability, but it requires coordinating routing, state, and capacity across locations.
Active-passive
Active-passive: a hot standby cluster or region stays ready but idle, taking over if the primary fails.
Structurally simpler, but you have to weigh standby utilization against failover speed.
Autoscaling
Autoscaling: automatically adding or removing model replicas based on traffic or utilization.
The goal is to match capacity to demand while holding latency SLAs and cutting waste.
Autoscaling window
Autoscaling window: the rolling time range that triggers a scale-up or scale-down.
A longer window is more stable; a shorter one reacts faster to sudden spikes.
Cold start
Cold start: the time from scaling a replica up from zero to its first successful response, including GPU provisioning, container startup, model loading, and engine compilation.
It's the main cost of scale-to-zero and directly affects queueing and first-request latency.
Scale to zero
Scale to zero: shut down all replicas when idle and spin them up on demand when requests arrive.
Saves cost, but depends on fast cold starts and reliable queueing.
Blue-green deployment
Blue-green deployment: maintain two parallel production environments and shift traffic between blue and green.
Enables zero-downtime releases and fast rollbacks.
Canary deployment
Canary deployment: send a small share of live traffic to the new version first, watch its stability and performance, then ramp up gradually.
A way to validate progress on real traffic.
Shadow traffic
Shadow traffic: mirror real production requests to a candidate deployment, but don't use the candidate's results to serve users.
Lets you evaluate a new version without affecting the main path.
Routing (inference)
Inference routing: send requests to the right replica based on load, KV cache, available LoRAs, and sequence shape.
A good router does more than round-robin — it also uses the model's runtime state.
gRPC
gRPC: a structured, schema-first, bidirectional streaming protocol.
Good for well-defined service-to-service communication; schema validation adds some overhead.
WebSocket
WebSocket: a lightweight, bidirectional streaming protocol.
Well suited to unstructured audio chunks and real-time user experience.
Service Level Agreement (SLA)
SLA: a contractual commitment the system makes about metrics such as latency, throughput, and availability.
An outward-facing promise, usually backed by observability data.
Service Level Objective (SLO)
SLO: an internal target the system sets to meet or exceed an SLA.
Internal targets should generally leave headroom above the external SLA.
Docker
Docker: packages an inference service and its dependencies into a standardized runtime unit using containers.
Makes deployment environments more reproducible, but doesn't automatically solve GPU driver or performance issues.
Dockerfile
Dockerfile: a human-readable, machine-executable instruction file for building a container image.
Pins down runtime dependencies, startup commands, and the environment.
8. Inference Engines, Runtimes, and Engineering Tools
Term
What it means
What to look at in an inference system
vLLM
A widely adopted inference engine with broad model and hardware support and mature defaults.
Common capabilities include continuous batching, KV management, quantization, and serving APIs.
SGLang
A fast inference engine with a flexible frontend/backend and strengthened MoE support.
Suited to scenarios that need structured workflows and high-performance serving.
TensorRT
NVIDIA's runtime optimized for high-performance inference.
Combines optimizations such as fused kernels and quantization.
TensorRT-LLM
An LLM inference engine built by NVIDIA, offering a Python API and supporting both TensorRT engines and a PyTorch backend.
Integrates fused kernels, quantization, and speculative decoding.
Triton Inference Server
NVIDIA's production-grade serving framework, supporting multiple backends.
It handles model serving orchestration, protocols, and backend integration — it isn't a single kernel library.
ONNX
An intermediate representation for models and its runtime ecosystem.
Used to exchange model representations across different frameworks or runtimes.
NIM
Pre-packaged, containerized microservices that NVIDIA provides for specific models.
Lowers the barrier to packaging model serving, but you still need to watch the underlying hardware and runtime behavior.
NVIDIA Dynamo
NVIDIA's open-source distributed serving platform, aimed at KV reuse, disaggregated inference, and multi-GPU/multi-node orchestration.
It sits at the serving orchestration layer, connecting model execution and cluster resources.
Transformers (library)
A reference implementation library for LLMs and other Transformer models.
Commonly used for model loading, configuration, inference logic, and research prototypes.
Diffusers (library)
A reference implementation library for image and video generation pipelines.
Strings components like text encoding, denoising, and the VAE into a generation pipeline.
ComfyUI
A workflow tool for assembling image generation pipelines.
Uses nodes to combine components such as the base model, refiner, LoRA, and ControlNet.
9. Diffusion, Multimodal, Vision, and Speech Inference
Term
What it means
What to look at in an inference system
Image generation pipeline
Image generation pipeline: typically composed of several models, such as a text encoder, an iterative denoiser, and a VAE.
It's not as simple as one forward pass of one model — multiple stages run in sequence.
Denoising model
Denoising model: the core of a diffusion pipeline, repeatedly refining latent noise into an image or video.
Inference steps, latent size, and kernels directly affect latency.
Iterative denoising (diffusion)
Iterative denoising: starting from noise, generating an image or video step by step in latent space.
Each step runs denoising computation, and the number of steps is a major performance lever.
Latent space (images/videos)
A low-dimensional latent space for images/videos where denoising usually happens, for example a 128x128 representation.
Computing in latent space uses fewer resources than computing directly on pixels.
VAE (variational autoencoder)
Variational autoencoder: decodes latents into pixel space; during training it can also encode pixels back into latents.
It's the key boundary in a diffusion pipeline between the latent and the final image.
CLIP (text encoder)
The CLIP text/image encoder, commonly used in early image pipelines such as SDXL.
Modern systems often replace or augment prompt understanding with a stronger, full LLM.
Classifier-free guidance
Classifier-free guidance: at each step, balancing unconditional and prompt-conditioned denoising.
Lower guidance is more creative; higher guidance follows the prompt more closely, but can affect quality and speed.
Few-step image generation
Few-step image generation: producing a usable image in eight steps or fewer.
Can be 80–90% faster, but usually comes with a clear quality trade-off — good for real-time scenarios.
Latent consistency
Latent consistency: a few-step strategy that directly predicts the target latent and can be refined repeatedly.
Very fast, but fidelity is usually lower than full diffusion.
SDXL
A representative early diffusion image pipeline: base + refiner + CLIP.
Modern systems usually keep the multi-stage structure but swap in stronger components.
Omni-modal
Omni-modal: accepts many kinds of input — text, images, video, audio — and produces many kinds of output too.
The key point is that neither input nor output modalities are limited to text.
Vision-language model (VLM)
Vision-language model: takes images/video and a text prompt as input and outputs text.
Inference has to account for visual encoding, cross-modal attention, and text decoding.
Encoder
Encoder: a network that turns raw input into an internal representation, such as Whisper's audio feature encoder.
Paired with a decoder in encoder-decoder models, it can serve as an early stage in a multimodal pipeline.
Automatic Speech Recognition (ASR)
Automatic Speech Recognition: a transcription model that takes audio input and produces text output, such as Whisper.
Decoder work often dominates runtime and can benefit from LLM-style optimizations and in-flight batching.
Voice activity detection (VAD)
Voice activity detection: splits an audio stream or file into segments containing speech.
A lightweight model used in ASR preprocessing to cut wasted compute on segments with no speech.
Diarization
Diarization: answers 'who spoke when', splitting audio by speaker.
Often paired with VAD to add speaker boundaries to ASR output.
Neural audio codec
Neural audio codec: compresses audio into tokens, which a matching decoder turns back into audio.
Turns audio generation into a token-level modeling problem.
SNAC (audio decoder)
A high-performance audio decoder path, often used alongside TTS token streams.
Its job is to turn a stream of audio tokens back into playable audio.
10. Speculative Decoding and Sampling Acceleration
These techniques all share one goal: make a single expensive forward pass through the target model confirm as many tokens as possible.
Term
What it means
What to watch in an inference system
Speculative decoding
Speculative decoding: draft tokens are generated first, then verified in a batch by the target model so a single forward pass can accept several tokens at once.
The key metrics are draft cost, acceptance rate, and how many tokens end up accepted per step.
EAGLE (speculation)
EAGLE: a small, specially trained draft model that consumes hidden states and proposes multiple tokens.
The aim is to make the draft align more closely with the target, raising the acceptance rate.
Medusa (speculation)
Medusa: adds extra decoder heads through fine-tuning so each forward pass produces several draft tokens.
It doesn't necessarily need a separate draft model, but it does require extra training or changes to the model architecture.
Lookahead decoding
Lookahead decoding: builds n-grams during inference so draft tokens can be predicted without a separate model.
It exploits patterns already present in the generation to cut down on serial dependencies.
N-gram speculation
N-gram speculation: uses n-grams observed during prefill to propose longer draft sequences at decode time.
Especially effective for code completion, since code tends to repeat itself and follow stable local patterns.
11. Retrieval, Embeddings, and Model Quality Evaluation
Term
What it means
What to watch in an inference system
Embedding model
Embedding model: encodes text or images into fixed-dimension vectors for semantic similarity.
Commonly powers RAG, search, and agent memory; modern versions often use an LLM backbone.
Matryoshka representations (embeddings)
Matryoshka representations: nested vector representations where the front of the vector carries more semantic weight, so dimensions can be truncated as needed.
Offers a tunable trade-off between vector size, retrieval cost, and quality.
Vector database
Vector database: stores and queries the semantic vectors produced by embeddings.
The retrieval layer of RAG; performance hinges on indexing, recall, filtering, and storage.
Vector similarity
Vector similarity: measures how close two vectors are using formulas such as cosine similarity.
Similar vectors usually mean similar semantics, but similarity is not factual correctness.
Retrieval-augmented generation (RAG)
Retrieval-augmented generation: an application pattern that fetches extra context for an LLM beyond the prompt itself.
End-to-end latency must include retrieval, reranking, context assembly, and the extra prefill.
Benchmark (intelligence)
Intelligence benchmark: measures how well a model answers questions or takes appropriate actions, e.g. MMLU.
It gauges model capability, not actual product quality.
Benchmark (performance)
Performance benchmark: measures the latency and throughput of an inference service under a given model and workload.
Hardware, ISL/OSL, concurrency, batch size, and measurement methodology must all be fixed.
Evals
Evals: task-specific tests that simulate real use cases to measure model intelligence in production scenarios.
Closer to the business than generic benchmarks, but they demand more work to maintain the data and criteria yourself.
Elo (quality meta-metric)
Elo as a quality meta-metric: compares model quality by pairwise win rates.
A directional signal only; it shouldn't replace task-specific metrics or human review on its own.
Goodhart's Law
Goodhart's Law: once a measure becomes a target, it stops being a good measure.
When optimizing TPS, TTFT, or a benchmark, keep an eye on quality, cost, and the real user experience at the same time.
12. GPU Architectures, Chips, and System Models
This section collects hardware terms that come with lots of names but shouldn't be confused with algorithmic concepts. The figures and roadmaps here are a snapshot of the book's edition.
Term
What it means
What to watch in an inference system
Ada Lovelace (architecture)
NVIDIA's graphics-oriented GPU architecture, released around the same time as Hopper; suited to small models and cost-sensitive workloads.
The book considers it a poor fit for large-scale LLM inference.
Ampere (architecture)
An earlier NVIDIA GPU architecture, still used in legacy or small-scale deployments.
When comparing it with Hopper and Blackwell, focus on cost, bandwidth, and efficiency at scale.
Hopper (architecture)
NVIDIA's 2022 GPU architecture, supporting FP8 and asynchronous programming features.
One of the target architectures for FlashAttention 3.
Blackwell (architecture)
NVIDIA's GPU generation from late 2024, supporting FP4, MXFP8, MXFP4, NVFP4, and high memory bandwidth.
In an inference context, the focus is on low-precision formats, bandwidth, and serving large models.
Rubin (architecture)
NVIDIA's next-generation architecture as described in the book, introducing HBM4 and CPX for compute-bound workloads.
A future/roadmap term as of the book's publication; check the latest official specs before relying on it.
Feynman (architecture)
A future NVIDIA architecture following Rubin, described in the book with limited detail.
Treat it as a roadmap name only; don't infer shipping capabilities from it.
B200
A Blackwell data center GPU; the book describes it as having 192 GB of VRAM, 8 TB/s of bandwidth, and 5 petaFLOPS of FP8.
These are the book's snapshot figures; actual SKUs and measurement methodology need to be verified separately.
B300
A Blackwell data center GPU; the book describes it as having 288 GB of VRAM, 8 TB/s of bandwidth, and 5 petaFLOPS of FP8.
Mainly useful for understanding how VRAM, bandwidth, and precision affect inference.
Grace CPU
NVIDIA's ARM CPU, working alongside GPUs over a high-bandwidth chip-to-chip interconnect.
Suited to systems that need fast CPU/GPU exchange, such as KV offload.
Vera CPU
An NVIDIA ARM CPU described in the book as succeeding Grace in the Rubin GPU generation.
This is a roadmap architecture name; specific product capabilities need to be verified per release.
GB200
An NVIDIA superchip combining a Grace CPU with a B200 GPU, linked chip-to-chip over high-bandwidth NVLink.
Suited to techniques that benefit from NVLink-C2C, such as KV cache offload and LoRA swapping.
GH200
An NVIDIA superchip combining a Grace CPU with an H200 GPU, linked chip-to-chip over high-bandwidth NVLink.
Similar to GB200, with the emphasis on high-bandwidth CPU/GPU cooperation.
NVL72
A rack-scale Blackwell system described in the book as interconnecting 72 GPUs and 36 CPUs.
Built for very large models and extremely high-throughput serving; system-level topology matters more than single-card specs.
References
This article compiles the terminology from Appendix A of Inference Engineering by Philip Kiely, published by Baseten Books. The original book orders terms alphabetically; this article regroups them along the inference system stack. For hardware models, software versions, and future architecture roadmaps, defer to the book's edition and official sources.