The stack separates execution responsibilities that are often collapsed into the single phrase “AI runtime.” Start at the layer that owns the constraint you are trying to change.
Hardware and System Substrate
Accelerators, memory hierarchy, drivers, interconnects, isolation, and system limits.
Kernels and Hardware Libraries
Tensor kernels, collectives, communication libraries, and device-specific execution paths.
Compiler and Graph Runtimes
Graph capture, IR, rewriting, partitioning, lowering, code generation, and memory planning.
Model and LLM Inference
Model loading, quantization, prefill, decode, KV cache, batching, streaming, and structured generation.
Serving and Distributed Execution
Network APIs, model repositories, scheduling, autoscaling, rollouts, and multi-host execution.
Edge and Browser Runtimes
Local, browser, mobile, NPU, offline, privacy-preserving, and constrained deployment.
Agentic and Application Runtimes
Identity, context, tools, memory, policy, approval, evaluation, evidence, and recovery.
Product and Workflow Layer
Domain records, user experience, business workflow, outcomes, and product-specific state.
