# ARuntime.com — Full LLM Reference Index Release: v3.9.0; reviewed 2026-06-27 UTC. ## Canonical positioning An AI runtime is the execution environment that turns model artifacts or model requests into operational behavior. Depending on the layer, it may compile computational graphs, schedule hardware, execute inference, serve models, coordinate distributed workloads, or govern context, tools, memory, policy, and traces. AI runtime is an umbrella term, not a single product category. Correct architecture begins by naming the runtime layer under discussion. The seven layers are: Layer 0 hardware and system substrate; Layer 1 kernels and hardware libraries; Layer 2 compiler and graph runtime; Layer 3 model and LLM inference engine; Layer 4 serving and distributed runtime; Layer 5 agentic and application runtime; Layer 6 product and workflow layer. An inference engine produces model outputs. An agentic application runtime produces controlled work. These are responsibilities at different layers, not substitute product categories. Categories describe responsibilities, not marketing labels. A product may implement several runtime layers. ## Primary navigation ### Start Landing page: https://aruntime.com/start-here/ - AI Runtime Overview — https://aruntime.com/overview/ - Taxonomy — https://aruntime.com/taxonomy/ - Seven-Layer Stack — https://aruntime.com/runtime-stack-diagram/ - How Runtimes Work — https://aruntime.com/how-ai-runtimes-work/ - Selection Guide — https://aruntime.com/runtime-selection-guide/ - Glossary — https://aruntime.com/glossary/ ### Runtime Types Landing page: https://aruntime.com/runtime-types/ - Emerging Runtime Types — https://aruntime.com/runtime-types/ - Neural Execution Engine — https://aruntime.com/runtime-types/neural-execution-engine/ - Autonomous Inference Runtime — https://aruntime.com/runtime-types/autonomous-inference-runtime/ - Agentic Execution Environment — https://aruntime.com/runtime-types/agentic-execution-environment/ - Machine Intelligence Runtime — https://aruntime.com/runtime-types/machine-intelligence-runtime/ - Semantic Runtime Engine — https://aruntime.com/runtime-types/semantic-runtime-engine/ - Cognitive Runtime Environment — https://aruntime.com/runtime-types/cognitive-runtime-environment/ ### Build & Operate Landing page: https://aruntime.com/build-operate/ - Reference Architecture — https://aruntime.com/reference-architecture/ - Runtime Architecture — https://aruntime.com/runtime-architecture/ - Control and Execution Planes — https://aruntime.com/control-execution-planes/ - Deployment Patterns — https://aruntime.com/deployment-patterns/ - Model Serving — https://aruntime.com/model-serving-orchestration/ - Distributed Inference — https://aruntime.com/distributed-inference/ - Edge and Browser — https://aruntime.com/edge-browser-runtimes/ - Agentic Runtimes — https://aruntime.com/agentic-runtimes/ - Security and Governance — https://aruntime.com/security-governance/ - Observability — https://aruntime.com/runtime-observability/ - Benchmarking — https://aruntime.com/benchmarking/ - Cost and Capacity — https://aruntime.com/cost-capacity-planning/ - Failure Recovery — https://aruntime.com/failure-recovery/ ### Developers Landing page: https://aruntime.com/developers/ - Developer Guide — https://aruntime.com/developer-guide/ - Examples — https://aruntime.com/examples/ - Integration Checklists — https://aruntime.com/integration-checklists/ - Runtime Request — https://aruntime.com/runtime-request-contract/ - Tool Contract — https://aruntime.com/tool-contract/ - Policy Approval — https://aruntime.com/policy-approval-contract/ - Evidence Schema — https://aruntime.com/evidence-schema/ - Trace Schema — https://aruntime.com/trace-schema/ ### Directory Landing page: https://aruntime.com/runtimes/ - Runtime Directory — https://aruntime.com/runtimes/ - Compare Categories — https://aruntime.com/compare/ - Comparison Scope — https://aruntime.com/product-comparison-scope/ - Directory Methodology — https://aruntime.com/directory-methodology/ - Corrections — https://aruntime.com/corrections/ ### Trust Landing page: https://aruntime.com/research/ - Reports — https://aruntime.com/reports/ - Protocols and Ecosystem — https://aruntime.com/protocols-ecosystem/ - Industry Landscape — https://aruntime.com/industry-landscape/ - Resources — https://aruntime.com/resources/ - Content Quality and Discovery — https://aruntime.com/ai-ready-web/ - Research Methodology — https://aruntime.com/research-methodology/ - Editorial Policy — https://aruntime.com/editorial-policy/ - Changelog — https://aruntime.com/changelog/ - Contact — https://aruntime.com/contact/ ## Managed page catalog ### Home URL: https://aruntime.com/ Status: Foundational; category: Start Here. Audience: Readers new to AI runtimes; platform architects; technical decision-makers Search intent: AI runtime definition and stack orientation Summary: ARuntime.com explains artificial intelligence runtimes across the full execution stack. Last reviewed: 2026-06-23 UTC. ### Start Here URL: https://aruntime.com/start-here/ Status: Foundational; category: Start Here. Audience: Readers new to AI runtimes and role-based learners Search intent: Orient readers to the AI runtime stack Summary: Start with definitions, the layered runtime stack, execution paths, use cases, and a glossary. Last reviewed: 2026-06-23 UTC. ### AI Runtime Overview URL: https://aruntime.com/overview/ Status: Foundational; category: Start Here. Audience: Readers who need a precise definition; architects; engineering leaders Search intent: Define AI runtime and distinguish its layers Summary: A vendor-neutral overview of AI runtime layers, from compilers and inference engines to model serving, deployment, agentic execution, security, and governance. Last reviewed: 2026-06-23 UTC. ### AI Runtime Taxonomy URL: https://aruntime.com/taxonomy/ Status: Foundational; category: Start Here. Audience: Platform architects; ML systems engineers; technical analysts Search intent: Classify AI runtime categories and boundaries Summary: A vendor-neutral taxonomy of AI runtime layers, categories, execution models, deployment boundaries, hardware targets, and system responsibilities. Last reviewed: 2026-06-23 UTC. ### How AI Runtimes Work URL: https://aruntime.com/how-ai-runtimes-work/ Status: Architecture; category: Start Here. Audience: Infrastructure engineers; application developers; technical decision-makers Search intent: Explain model and request execution paths Summary: How model artifacts become executable programs and how production requests move through scheduling, inference, tools, policy, telemetry, memory, and failure handling. Last reviewed: 2026-06-23 UTC. ### Runtime Stack Diagram URL: https://aruntime.com/runtime-stack-diagram/ Status: Foundational; category: Start Here. Audience: All readers Search intent: View an accessible AI runtime stack diagram Summary: An accessible diagram and text description of the seven-layer AI runtime stack. Last reviewed: 2026-06-23 UTC. ### AI Runtime Glossary URL: https://aruntime.com/glossary/ Status: Foundational; category: Foundations. Audience: Readers learning AI runtime concepts; technical decision-makers Search intent: A searchable glossary of AI runtime terminology across compilers, inference engines, serving systems, edge/browser runtimes, agentic runtimes, security, and observability. Summary: Definitions for the complete AI runtime stack, generated from the reviewed version-controlled glossary catalog. Last reviewed: 2026-06-23 UTC. ### Runtime Stack URL: https://aruntime.com/runtime-stack/ Status: Foundational; category: Runtime Stack. Audience: Platform architects; AI infrastructure engineers; ML systems engineers Search intent: Navigate the seven AI runtime layers Summary: Navigate hardware, kernels, compilers, inference engines, serving, agentic runtimes, and product workflows. Last reviewed: 2026-06-23 UTC. ### Hardware and System Substrate URL: https://aruntime.com/hardware-system-substrate/ Status: Architecture; category: Runtime Stack. Audience: Infrastructure and ML systems engineers Search intent: Explain hardware substrate responsibilities in AI runtime architecture Summary: Hardware, memory, drivers, interconnects, isolation, topology, power, and system constraints beneath AI runtimes. Last reviewed: 2026-06-23 UTC. ### Kernels and Hardware Libraries URL: https://aruntime.com/kernels-hardware-libraries/ Status: Architecture; category: Runtime Stack. Audience: ML systems engineers; compiler and performance engineers Search intent: Explain AI kernel and hardware-library runtime responsibilities Summary: Tensor kernels, device libraries, collectives, communication paths, and low-level performance boundaries. Last reviewed: 2026-06-23 UTC. ### Execution Models URL: https://aruntime.com/execution-models/ Status: Foundational; category: Runtime Stack. Audience: Readers learning AI runtime concepts; technical decision-makers Search intent: Compare eager, graph, JIT, AOT, interpreted, dataflow, and hybrid AI execution models, including dynamic shapes, guards, warmup, caches, fallback, and deployment trade-offs. Summary: Compare eager, graph, JIT, AOT, interpreted, dataflow, and hybrid AI execution models, including dynamic shapes, guards, warmup, caches, fallback, and deployment trade-offs. Last reviewed: 2026-06-23 UTC. ### Hardware Targets URL: https://aruntime.com/hardware-targets/ Status: Foundational; category: Runtime Stack. Audience: Readers learning AI runtime concepts; technical decision-makers Search intent: Understand CPU, GPU, TPU, NPU, FPGA, browser, mobile, and heterogeneous AI hardware targets through memory bandwidth, precision, interconnect, power, and runtime constraints. Summary: Understand CPU, GPU, TPU, NPU, FPGA, browser, mobile, and heterogeneous AI hardware targets through memory bandwidth, precision, interconnect, power, and runtime constraints. Last reviewed: 2026-06-23 UTC. ### Model Formats and Intermediate Representations URL: https://aruntime.com/model-formats-and-irs/ Status: Foundational; category: Runtime Stack. Audience: Readers learning AI runtime concepts; technical decision-makers Search intent: Compare ONNX, StableHLO, MLIR dialects, exported PyTorch programs, GGUF, TensorRT engines, ExecuTorch PTE files, and auxiliary model assets by role and runtime implications. Summary: Compare ONNX, StableHLO, MLIR dialects, exported PyTorch programs, GGUF, TensorRT engines, ExecuTorch PTE files, and auxiliary model assets by role and runtime implications. Last reviewed: 2026-06-23 UTC. ### Compiler and Graph Runtimes URL: https://aruntime.com/compiler-pipeline/ Status: Architecture; category: Runtime Stack. Audience: ML systems engineers; AI infrastructure engineers Search intent: A practical AI compiler pipeline guide covering graph capture, intermediate representations, optimization, partitioning, lowering, code generation, memory planning, AOT/JIT, fallback, and warmup. Summary: A practical AI compiler pipeline guide covering graph capture, intermediate representations, optimization, partitioning, lowering, code generation, memory planning, AOT/JIT, fallback, and warmup. Last reviewed: 2026-06-23 UTC. ### Optimization Techniques URL: https://aruntime.com/optimization-techniques/ Status: Architecture; category: Runtime Stack. Audience: ML systems engineers; AI infrastructure engineers Search intent: A practical guide to graph rewrites, fusion, tiling, memory planning, FlashAttention-style kernels, quantization, sparsity, batching, caching, speculation, and autotuning. Summary: A practical guide to graph rewrites, fusion, tiling, memory planning, FlashAttention-style kernels, quantization, sparsity, batching, caching, speculation, and autotuning. Last reviewed: 2026-06-23 UTC. ### Model and LLM Inference URL: https://aruntime.com/llm-inference/ Status: Architecture; category: Runtime Stack. Audience: ML systems engineers; AI infrastructure engineers Search intent: Understand LLM inference mechanics, including prefill, decode, KV cache, continuous batching, prefix reuse, quantization, structured generation, cancellation, tail latency, and production metrics. Summary: Understand LLM inference mechanics, including prefill, decode, KV cache, continuous batching, prefix reuse, quantization, structured generation, cancellation, tail latency, and production metrics. Last reviewed: 2026-06-23 UTC. ### KV Cache and Runtime Memory URL: https://aruntime.com/kv-cache-memory/ Status: Architecture; category: Runtime Stack. Audience: ML systems engineers; AI infrastructure engineers Search intent: A deep guide to KV-cache architecture, PagedAttention, prefix caching, RadixAttention, tiered offload, cache-aware routing, isolation, memory sizing, and observability. Summary: A deep guide to KV-cache architecture, PagedAttention, prefix caching, RadixAttention, tiered offload, cache-aware routing, isolation, memory sizing, and observability. Last reviewed: 2026-06-23 UTC. ### Scheduling and Batching URL: https://aruntime.com/scheduling-and-batching/ Status: Architecture; category: Runtime Stack. Audience: ML systems engineers; AI infrastructure engineers Search intent: Guide to AI inference scheduling and batching: admission control, continuous batching, dynamic batching, chunked prefill, fairness, priorities, backpressure, cancellation, and Goodput. Summary: Guide to AI inference scheduling and batching: admission control, continuous batching, dynamic batching, chunked prefill, fairness, priorities, backpressure, cancellation, and Goodput. Last reviewed: 2026-06-23 UTC. ### Speculative Decoding URL: https://aruntime.com/speculative-decoding/ Status: Architecture; category: Runtime Stack. Audience: ML systems engineers; AI infrastructure engineers Search intent: Learn speculative decoding mechanics, draft and target models, verification, acceptance rate, EAGLE, Medusa, multi-token prediction, integration constraints, metrics, and failure modes. Summary: Learn speculative decoding mechanics, draft and target models, verification, acceptance rate, EAGLE, Medusa, multi-token prediction, integration constraints, metrics, and failure modes. Last reviewed: 2026-06-23 UTC. ### Model Serving URL: https://aruntime.com/model-serving-orchestration/ Status: Architecture; category: Runtime Stack. Audience: Platform architects; infrastructure engineers Search intent: A practical guide to model servers and serving platforms: APIs, repositories, loading, readiness, batching, multi-model hosting, autoscaling, canaries, rollback, backpressure, and observability. Summary: A practical guide to model servers and serving platforms: APIs, repositories, loading, readiness, batching, multi-model hosting, autoscaling, canaries, rollback, backpressure, and observability. Last reviewed: 2026-06-23 UTC. ### Distributed Inference URL: https://aruntime.com/distributed-inference/ Status: Architecture; category: Runtime Stack. Audience: Platform architects; infrastructure engineers Search intent: Guide to tensor, pipeline, data, expert, context, sequence, and disaggregated prefill/decode inference, including interconnects, KV transfer, locality, scheduling, and failure handling. Summary: Guide to tensor, pipeline, data, expert, context, sequence, and disaggregated prefill/decode inference, including interconnects, KV transfer, locality, scheduling, and failure handling. Last reviewed: 2026-06-23 UTC. ### Edge and Browser Runtimes URL: https://aruntime.com/edge-browser-runtimes/ Status: Architecture; category: Runtime Stack. Audience: Browser, mobile, edge, and privacy-focused developers Search intent: Compare edge and browser AI runtime boundaries Summary: Local, browser, mobile, NPU, offline, constrained, and privacy-preserving AI runtime deployment. Last reviewed: 2026-06-23 UTC. ### Browser Runtimes URL: https://aruntime.com/browser-runtimes/ Status: Production guidance; category: Runtime Stack. Audience: Infrastructure engineers; developers selecting deployment targets Search intent: Build browser AI runtimes with WebAssembly, WebGPU, WebNN, ONNX Runtime Web, Workers, model caching, I/O binding, graph capture, progressive enhancement, privacy, and fallback. Summary: Build browser AI runtimes with WebAssembly, WebGPU, WebNN, ONNX Runtime Web, Workers, model caching, I/O binding, graph capture, progressive enhancement, privacy, and fallback. Last reviewed: 2026-06-23 UTC. ### Edge, Mobile, and TinyML URL: https://aruntime.com/edge-mobile-tinyml/ Status: Production guidance; category: Runtime Stack. Audience: Infrastructure engineers; developers selecting deployment targets Search intent: Design edge, mobile, and TinyML runtimes with AOT artifacts, ExecuTorch, delegates, quantization, CPU/GPU/NPU partitioning, offline operation, fleet updates, and real-time constraints. Summary: Design edge, mobile, and TinyML runtimes with AOT artifacts, ExecuTorch, delegates, quantization, CPU/GPU/NPU partitioning, offline operation, fleet updates, and real-time constraints. Last reviewed: 2026-06-23 UTC. ### Local Desktop Runtimes URL: https://aruntime.com/local-desktop/ Status: Production guidance; category: Runtime Stack. Audience: Infrastructure engineers; developers selecting deployment targets Search intent: Guide to local AI runtimes on desktops and workstations, including quantized packages, CPU/GPU offload, local model servers, memory mapping, privacy, updates, and benchmarking. Summary: Guide to local AI runtimes on desktops and workstations, including quantized packages, CPU/GPU offload, local model servers, memory mapping, privacy, updates, and benchmarking. Last reviewed: 2026-06-23 UTC. ### Agentic and Application Runtimes URL: https://aruntime.com/agentic-runtimes/ Status: Production guidance; category: Runtime Stack. Audience: Agent developers; platform architects; governance leaders Search intent: Define the agentic application-runtime layer Summary: A production guide to agentic runtimes: identity, context, MCP tools and resources, typed execution, memory scopes, durable workflows, human approval, policy, replay, evaluation, and failure containment. Last reviewed: 2026-06-23 UTC. ### Product and Workflow Layer URL: https://aruntime.com/product-workflow-layer/ Status: Architecture; category: Runtime Stack. Audience: Application architects; product engineers; technical product leaders Search intent: Explain the product layer above AI runtimes Summary: User experience, domain records, business workflow, product state, and outcomes above runtime execution. Last reviewed: 2026-06-23 UTC. ### Build and Operate URL: https://aruntime.com/build-operate/ Status: Production guidance; category: Build and Operate. Audience: Platform architects; operators; security and governance leaders Search intent: Design and operate production AI runtime architecture Summary: Design runtime planes, deployment topologies, security, observability, evaluation, capacity, and recovery. Last reviewed: 2026-06-23 UTC. ### Reference Architecture URL: https://aruntime.com/reference-architecture/ Status: Production guidance; category: Build and Operate. Audience: Platform architects; implementation teams; security and governance leaders Search intent: Provide a concrete AI runtime implementation blueprint Summary: Detailed AI runtime reference architecture with request gateway, identity, context providers, model router, inference adapters, tool broker, memory, policy, workflow, evaluation, telemetry, and deployment variants. Last reviewed: 2026-06-23 UTC. ### Runtime Architecture URL: https://aruntime.com/runtime-architecture/ Status: Architecture; category: Build and Operate. Audience: Platform architects; infrastructure engineers; security architects Search intent: Explain AI runtime architecture concepts and boundaries Summary: A four-plane AI runtime reference architecture covering control, context, execution, trust, component contracts, deployment topologies, failure domains, and portability. Last reviewed: 2026-06-23 UTC. ### Control and Execution Planes URL: https://aruntime.com/control-execution-planes/ Status: Architecture; category: Build and Operate. Audience: Platform architects; operators; security engineers Search intent: Separate AI runtime control and execution planes Summary: Control, execution, data, and evidence-plane ownership, data flow, trust boundaries, and degraded behavior. Last reviewed: 2026-06-23 UTC. ### Deployment Patterns URL: https://aruntime.com/deployment-patterns/ Status: Production guidance; category: Build and Operate. Audience: Infrastructure engineers; developers selecting deployment targets Search intent: Compare browser, local, mobile, edge, private-cloud, managed-cloud, Kubernetes, serverless, hybrid, air-gapped, and benchmark AI runtime deployment patterns. Summary: Compare browser, local, mobile, edge, private-cloud, managed-cloud, Kubernetes, serverless, hybrid, air-gapped, and benchmark AI runtime deployment patterns. Last reviewed: 2026-06-23 UTC. ### Cloud and Data Center Runtimes URL: https://aruntime.com/cloud-data-center/ Status: Production guidance; category: Build and Operate. Audience: Infrastructure engineers; developers selecting deployment targets Search intent: Design cloud and private data-center AI runtimes across GPU clusters, Kubernetes, managed platforms, autoscaling, networking, tenancy, resilience, and cost. Summary: Design cloud and private data-center AI runtimes across GPU clusters, Kubernetes, managed platforms, autoscaling, networking, tenancy, resilience, and cost. Last reviewed: 2026-06-23 UTC. ### Serverless AI Runtime Patterns URL: https://aruntime.com/serverless-runtime/ Status: Production guidance; category: Build and Operate. Audience: Infrastructure engineers; developers selecting deployment targets Search intent: Design serverless and microVM AI runtime patterns with cold-start budgets, model packaging, warm pools, scale-to-zero, burst admission, storage, observability, and side-effect safety. Summary: Design serverless and microVM AI runtime patterns with cold-start budgets, model packaging, warm pools, scale-to-zero, burst admission, storage, observability, and side-effect safety. Last reviewed: 2026-06-23 UTC. ### Hybrid AI Runtime Design URL: https://aruntime.com/hybrid-runtime/ Status: Production guidance; category: Build and Operate. Audience: Infrastructure engineers; developers selecting deployment targets Search intent: Design local-edge-cloud hybrid AI runtimes with policy routing, fallback, state reconciliation, model parity, cache boundaries, observability, and data residency. Summary: Design local-edge-cloud hybrid AI runtimes with policy routing, fallback, state reconciliation, model parity, cache boundaries, observability, and data residency. Last reviewed: 2026-06-23 UTC. ### Security and Governance URL: https://aruntime.com/security-governance/ Status: Production guidance; category: Build and Operate. Audience: Platform operators; security and governance leaders; engineering managers Search intent: AI runtime security and governance guide covering prompt injection, tool authorization, data exfiltration, model supply chain, sandboxing, memory poisoning, multi-tenancy, audit, approval, and incident response. Summary: AI runtime security and governance guide covering prompt injection, tool authorization, data exfiltration, model supply chain, sandboxing, memory poisoning, multi-tenancy, audit, approval, and incident response. Last reviewed: 2026-06-23 UTC. ### Runtime Observability URL: https://aruntime.com/runtime-observability/ Status: Architecture; category: Build and Operate. Audience: Platform architects; infrastructure engineers Search intent: Practical AI runtime observability covering traces, metrics, logs, token timing, cache metrics, tool and policy events, cost attribution, evaluation, replay, profiling, and privacy. Summary: Practical AI runtime observability covering traces, metrics, logs, token timing, cache metrics, tool and policy events, cost attribution, evaluation, replay, profiling, and privacy. Last reviewed: 2026-06-23 UTC. ### Evaluation URL: https://aruntime.com/evaluation/ Status: Production guidance; category: Build and Operate. Audience: ML engineers; platform teams; product and governance leaders Search intent: Evaluate AI runtime quality, safety, policy, and outcomes Summary: Versioned offline and online evaluation for model behavior, runtime controls, tool use, policy, and business outcomes. Last reviewed: 2026-06-23 UTC. ### Benchmarking AI Runtimes URL: https://aruntime.com/benchmarking/ Status: Production guidance; category: Build and Operate. Audience: Platform operators; security and governance leaders; engineering managers Search intent: Benchmark AI runtimes correctly with TTFT, TPOT, ITL, Goodput, throughput, tail latency, quality, concurrency, cache state, hardware, methodology, cost, power, and reproducibility. Summary: Benchmark AI runtimes correctly with TTFT, TPOT, ITL, Goodput, throughput, tail latency, quality, concurrency, cache state, hardware, methodology, cost, power, and reproducibility. Last reviewed: 2026-06-23 UTC. ### Cost and Capacity Planning URL: https://aruntime.com/cost-capacity-planning/ Status: Production guidance; category: Build and Operate. Audience: Platform operators; security and governance leaders; engineering managers Search intent: Plan AI runtime capacity and cost using model memory, KV cache, concurrency, Goodput, arrival rates, queueing, accelerator utilization, warm pools, provider cost, and task outcomes. Summary: Plan AI runtime capacity and cost using model memory, KV cache, concurrency, Goodput, arrival rates, queueing, accelerator utilization, warm pools, provider cost, and task outcomes. Last reviewed: 2026-06-23 UTC. ### Runtime Selection Guide URL: https://aruntime.com/runtime-selection-guide/ Status: Production guidance; category: Build and Operate. Audience: Platform operators; security and governance leaders; engineering managers Search intent: Choose an AI runtime using model format, latency, throughput, context, hardware, deployment, privacy, tools, durability, observability, licensing, maturity, and operational capability. Summary: Choose an AI runtime using model format, latency, throughput, context, hardware, deployment, privacy, tools, durability, observability, licensing, maturity, and operational capability. Last reviewed: 2026-06-23 UTC. ### Failure Recovery URL: https://aruntime.com/failure-recovery/ Status: Production guidance; category: Build and Operate. Audience: Platform engineers; application developers; incident responders Search intent: Design AI runtime failure detection and recovery Summary: Detection, retry eligibility, idempotency, rollback, compensation, escalation, and evidence for runtime failures. Last reviewed: 2026-06-23 UTC. ### Developers URL: https://aruntime.com/developers/ Status: Production guidance; category: Developers. Audience: Application developers; platform engineers; API and security engineers Search intent: Implement AI runtime contracts and governed execution Summary: Implement versioned request, tool, policy, evidence, and trace contracts with runnable examples. Last reviewed: 2026-06-23 UTC. ### Developer Guide URL: https://aruntime.com/developer-guide/ Status: Production guidance; category: Developers. Audience: Application developers; platform engineers Search intent: Build production AI runtimes with versioned request/response contracts, adapters, typed tools, routing, memory, policy checkpoints, traces, idempotency, streaming, evaluations, and SLOs. Summary: Build production AI runtimes with versioned request/response contracts, adapters, typed tools, routing, memory, policy checkpoints, traces, idempotency, streaming, evaluations, and SLOs. Last reviewed: 2026-06-23 UTC. ### Legacy Runtime Contract URL: https://aruntime.com/runtime-request-contract/ Status: Retired; category: Developers. Audience: Application developers; platform engineers; API designers Search intent: Legacy alias redirected to the canonical Runtime Request Contract. Summary: Retired in favor of the versioned Runtime Request Contract. Last reviewed: 2026-06-23 UTC. ### Runtime Request Contract URL: https://aruntime.com/runtime-request-contract/ Status: Production guidance; category: Developers. Audience: Application developers; platform and API engineers Search intent: Specify a versioned AI runtime request envelope Summary: A readable specification and JSON Schema for identity, tenancy, risk, permissions, route constraints, tools, memory, budgets, approvals, output, traces, and retention. Last reviewed: 2026-06-23 UTC. ### Tool Contract URL: https://aruntime.com/tool-contract/ Status: Production guidance; category: Developers. Audience: Agent developers; platform engineers; security engineers Search intent: Specify typed and governed AI tool execution Summary: Tool identity, schemas, authentication, permissions, side effects, egress, timeout, retry, idempotency, approval, compensation, audit, and deprecation. Last reviewed: 2026-06-23 UTC. ### Policy and Approval Contract URL: https://aruntime.com/policy-approval-contract/ Status: Production guidance; category: Developers. Audience: Security engineers; governance leaders; runtime developers Search intent: Specify policy decisions and approval gates for AI execution Summary: Versioned policy decisions, reason codes, approval authority, expiration, override rules, evidence, redaction, and review paths. Last reviewed: 2026-06-23 UTC. ### Evidence Schema URL: https://aruntime.com/evidence-schema/ Status: Production guidance; category: Developers. Audience: Platform engineers; auditors; governance and incident teams Search intent: Record minimized, correlated AI runtime evidence Summary: A minimized evidence record for request identity, versions, context references, decisions, tool effects, failures, retries, recovery, evaluation, redaction, and retention. Last reviewed: 2026-06-23 UTC. ### Trace Schema URL: https://aruntime.com/trace-schema/ Status: Production guidance; category: Developers. Audience: Platform engineers; SRE; observability and evaluation teams Search intent: Correlate infrastructure, model, tool, policy, outcome, and evaluation traces Summary: A trace envelope that separates infrastructure, model, tool, policy, business-outcome, and evaluation telemetry. Last reviewed: 2026-06-23 UTC. ### AI Runtime Examples URL: https://aruntime.com/examples/ Status: Production guidance; category: Developers. Audience: Application developers; platform engineers Search intent: Complete AI runtime architecture examples for local assistants, enterprise RAG, browser inference, mobile vision, high-throughput LLM serving, durable agents, and hybrid routing. Summary: Complete AI runtime architecture examples for local assistants, enterprise RAG, browser inference, mobile vision, high-throughput LLM serving, durable agents, and hybrid routing. Last reviewed: 2026-06-23 UTC. ### Integration Checklists URL: https://aruntime.com/integration-checklists/ Status: Production guidance; category: Developers. Audience: Application developers; platform engineers Search intent: Production AI runtime integration checklists for contracts, model adapters, tools, context/RAG, memory, security, observability, benchmarking, deployment, and operations. Summary: Production AI runtime integration checklists for contracts, model adapters, tools, context/RAG, memory, security, observability, benchmarking, deployment, and operations. Last reviewed: 2026-06-23 UTC. ### Runtime Directory URL: https://aruntime.com/runtimes/ Status: Research; category: Directory. Audience: Architects and developers evaluating runtime products Search intent: Browse structured AI runtime profiles Summary: Browse a vendor-neutral AI runtime directory covering compiler and graph runtimes, inference engines, model servers, serving platforms, edge, browser, and local runtime systems. Last reviewed: 2026-06-23 UTC. ### Compare Runtime Categories URL: https://aruntime.com/compare/ Status: Research; category: Directory. Audience: Technical evaluators; procurement teams; platform architects Search intent: Compare AI runtime categories within an explicit scope Summary: Compare AI runtimes responsibly by category, stack layer, model, hardware, precision, workload, SLO, quality, feature boundary, operations, licensing, and evidence. Last reviewed: 2026-06-23 UTC. ### Product Comparison Scope URL: https://aruntime.com/product-comparison-scope/ Status: Research; category: Directory. Audience: Technical evaluators and procurement teams Search intent: Define safeguards for AI runtime product comparisons Summary: Required scope, versions, hardware, workloads, exclusions, verification dates, and unverified fields for runtime comparisons. Last reviewed: 2026-06-23 UTC. ### Directory Methodology URL: https://aruntime.com/directory-methodology/ Status: Research; category: Directory. Audience: Vendors; maintainers; evaluators; readers Search intent: Explain runtime directory inclusion, categorization, and verification Summary: How runtime directory entries are sourced, categorized, versioned, reviewed, corrected, and marked when information is unverified. Last reviewed: 2026-06-23 UTC. ### Research URL: https://aruntime.com/research/ Status: Research; category: Research. Audience: Researchers; technical analysts; vendors; maintainers; readers Search intent: Navigate ARuntime research, sources, and editorial methods Summary: Review reports, protocols, source records, methodology, policy, corrections, and the changelog. Last reviewed: 2026-06-23 UTC. ### Reports URL: https://aruntime.com/reports/ Status: Research; category: Research. Audience: Researchers; architects; technical analysts Search intent: Browse long-form AI runtime reports Summary: Versioned reports on AI runtime architecture, execution models, deployment, governance, and ecosystem changes. Last reviewed: 2026-06-23 UTC. ### Protocols and Ecosystem URL: https://aruntime.com/protocols-ecosystem/ Status: Research; category: Research. Audience: Platform architects; protocol implementers; analysts Search intent: Map protocols that affect AI runtime interoperability Summary: Protocols and specifications for model formats, serving APIs, tools, identity, telemetry, policy, and evidence. Last reviewed: 2026-06-23 UTC. ### Industry Landscape URL: https://aruntime.com/industry-landscape/ Status: Research; category: Research. Audience: Technical decision-makers; analysts; engineering managers Search intent: Understand the AI runtime ecosystem by responsibility Summary: A responsibility-based landscape across compilers, inference engines, serving, distributed execution, edge, browser, and agentic application runtimes. Last reviewed: 2026-06-23 UTC. ### AI Runtime Resources URL: https://aruntime.com/resources/ Status: Foundational; category: Research. Audience: Technical readers Search intent: A curated resource hub for ARuntime.com reference pages, implementation guides, runtime directory entries, and trust resources. Summary: A curated resource hub for ARuntime.com reference pages, implementation guides, runtime directory entries, and trust resources. Last reviewed: 2026-06-23 UTC. ### Future Trends in AI Runtimes URL: https://aruntime.com/future-trends/ Status: Production guidance; category: Research. Audience: Platform operators; security and governance leaders; engineering managers Search intent: Evidence-based AI runtime trends: speculative decoding, disaggregated serving, cache-aware routing, AI PCs, WebNN, dynamic precision, energy-aware scheduling, confidential execution, and agent runtime infrastructure. Summary: Evidence-based AI runtime trends: speculative decoding, disaggregated serving, cache-aware routing, AI PCs, WebNN, dynamic precision, energy-aware scheduling, confidential execution, and agent runtime infrastructure. Last reviewed: 2026-06-23 UTC. ### Research Methodology URL: https://aruntime.com/research-methodology/ Status: Research; category: Research. Audience: Technical analysts; vendors; contributors Search intent: Explain research, source selection, and product verification Summary: ARuntime.com research methodology covering source discovery, primary-source verification, claim matrix, benchmark review, terminology, diagrams, code validation, freshness, and correction workflow. Last reviewed: 2026-06-23 UTC. ### Editorial Policy URL: https://aruntime.com/editorial-policy/ Status: Foundational; category: Research. Audience: Readers, vendors, contributors, and reviewers Search intent: Explain editorial standards and corrections Summary: ARuntime.com editorial policy for technical accuracy, vendor neutrality, source hierarchy, benchmark claims, time-sensitive facts, AI assistance, conflicts, corrections, and review dates. Last reviewed: 2026-06-23 UTC. ### Corrections URL: https://aruntime.com/corrections/ Status: Foundational; category: Research. Audience: Readers, vendors, maintainers, and contributors Search intent: Submit and review factual corrections Summary: Report corrections to ARuntime.com and review the process for verification, material corrections, source updates, security-sensitive reports, and publication records. Last reviewed: 2026-06-23 UTC. ### Content Quality and Discovery for ARuntime.com URL: https://aruntime.com/ai-ready-web/ Status: Production guidance; category: Research and Trust. Audience: Readers, maintainers, search crawlers, answer systems, and reviewers validating ARuntime public content quality Search intent: ARuntime content quality discovery files machine readable runtime pages boundaries Summary: Public record of how ARuntime.com keeps runtime pages visible, source-grounded, accessible, discoverable, and bounded without copied guidance or hidden bot text. Last reviewed: 2026-06-27 UTC. ### Answer-Ready Runtime Definitions URL: https://aruntime.com/ai-ready-web/answer-engine-optimization/ Status: Production guidance; category: Research and Trust. Audience: Editors, architects, and reviewers checking whether ARuntime runtime definitions can be summarized accurately Search intent: ARuntime answer ready runtime definitions direct answer canonical page Summary: How ARuntime.com structures its own runtime definitions and comparisons so answers remain concise, visible, citable, and bounded. Last reviewed: 2026-06-27 UTC. ### Generative Citation Support for Runtime Content URL: https://aruntime.com/ai-ready-web/generative-engine-optimization/ Status: Production guidance; category: Research and Trust. Audience: Maintainers, researchers, and reviewers checking generated summaries against ARuntime source pages Search intent: ARuntime generative citation support runtime content route inventory llms Summary: How ARuntime.com keeps runtime content retrievable, comparable, citable, and bounded for generative summaries without hidden prompts or copied guidance. Last reviewed: 2026-06-27 UTC. ### Search Discovery for AI Runtime Pages URL: https://aruntime.com/ai-ready-web/seo/ Status: Production guidance; category: Research and Trust. Audience: Site maintainers and technical reviewers auditing ARuntime search metadata and discovery files Search intent: ARuntime search discovery metadata sitemap robots structured data runtime pages Summary: Search-discovery implementation record for ARuntime.com titles, descriptions, canonical URLs, internal links, sitemap, robots, feeds, and visible structured data. Last reviewed: 2026-06-27 UTC. ### Changelog URL: https://aruntime.com/changelog/ Status: Foundational; category: Research. Audience: Readers, contributors, and maintainers Search intent: Review material site and taxonomy changes Summary: Material site, taxonomy, source, methodology, and correction changes with UTC dates. Last reviewed: 2026-06-27 UTC. ### About ARuntime.com URL: https://aruntime.com/about/ Status: Foundational; category: Utility. Audience: All readers Search intent: Explain site purpose and taxonomy ownership Summary: About ARuntime.com, its mission, layered AI runtime taxonomy, audience, editorial standards, Michael Kappel, and publication approach. Last reviewed: 2026-06-23 UTC. ### Contact URL: https://aruntime.com/contact/ Status: Foundational; category: Utility. Audience: Readers and professional contacts Search intent: Contact the site editor Summary: Contact Michael Kappel for AI runtime architecture, enterprise software modernization, technical collaboration, source recommendations, and ARuntime.com feedback. Last reviewed: 2026-06-23 UTC. ### Accessibility URL: https://aruntime.com/accessibility/ Status: Foundational; category: Utility. Audience: All visitors Search intent: Explain accessibility support and issue reporting Summary: Accessibility features, current conformance target, known limits, and issue-reporting route. Last reviewed: 2026-06-23 UTC. ### Privacy Policy URL: https://aruntime.com/privacy/ Status: Foundational; category: Legal. Audience: All visitors Search intent: Explain site privacy practices Summary: Privacy baseline for the ARuntime.com theme package and live-site considerations for production WordPress deployment. Last reviewed: 2026-06-23 UTC. ### Terms of Use URL: https://aruntime.com/terms/ Status: Foundational; category: Legal. Audience: All visitors Search intent: Explain site terms of use Summary: Terms of use for ARuntime.com as a public technical reference site. Last reviewed: 2026-06-23 UTC. ### Site Links and Navigation URL: https://aruntime.com/links/ Status: Retired; category: Trust. Audience: Readers, contributors, and vendors Search intent: Legacy link hub redirected to Start Here. Summary: Retired link hub. Requests redirect to the Start Here index. Last reviewed: 2026-06-23 UTC. ### Support and Corrections URL: https://aruntime.com/support/ Status: Retired; category: Trust. Audience: Readers, contributors, and vendors Search intent: Legacy support page redirected to the canonical corrections workflow. Summary: Retired support page. Requests redirect to the corrections workflow. Last reviewed: 2026-06-23 UTC. ### Emerging AI Runtime Types URL: https://aruntime.com/runtime-types/ Status: Emerging proposal; category: Architecture. Audience: Architects, AI platform engineers, technical decision makers Search intent: emerging AI runtime types AIR AEE MIR SRE CRE NEE Summary: A boundary-focused map of emerging AI runtime categories including AIR, AEE, MIR, SRE, CRE, and NEE. Last reviewed: 2026-06-27 UTC. ### Autonomous Inference Runtime URL: https://aruntime.com/runtime-types/autonomous-inference-runtime/ Status: Emerging proposal; category: Architecture. Audience: AI architects, platform engineers, runtime implementers Search intent: Autonomous Inference Runtime definition architecture runtime Summary: An Autonomous Inference Runtime is a local or edge-oriented execution environment that runs trained models close to the device, sensor, user, robot, vehicle, or industrial process that needs the result. Its primary job is to make inference fast, private, resilient, power-aware, a Last reviewed: 2026-06-27 UTC. ### Agentic Execution Environment URL: https://aruntime.com/runtime-types/agentic-execution-environment/ Status: Emerging proposal; category: Architecture. Audience: AI architects, platform engineers, runtime implementers Search intent: Agentic Execution Environment definition architecture runtime Summary: An Agentic Execution Environment is a runtime boundary that safely runs autonomous or semi-autonomous AI agent work. It provides sandbox isolation, resource limits, lifecycle controls, tool authorization, secret brokering, audit logging, policy hooks, and recovery for agent sessi Last reviewed: 2026-06-27 UTC. ### Machine Intelligence Runtime URL: https://aruntime.com/runtime-types/machine-intelligence-runtime/ Status: Emerging proposal; category: Architecture. Audience: AI architects, platform engineers, runtime implementers Search intent: Machine Intelligence Runtime definition architecture runtime Summary: A Machine Intelligence Runtime is an enterprise execution-time control layer that coordinates model-backed work across identity, policy, memory, tools, model routing, evaluation, telemetry, audit, budgets, and governed state changes. It separates probabilistic generation from det Last reviewed: 2026-06-27 UTC. ### Semantic Runtime Engine URL: https://aruntime.com/runtime-types/semantic-runtime-engine/ Status: Emerging proposal; category: Architecture. Audience: AI architects, platform engineers, runtime implementers Search intent: Semantic Runtime Engine definition architecture runtime Summary: A Semantic Runtime Engine is a runtime layer that manages meaning as an operational data type. It stores, retrieves, transforms, and governs semantic representations such as embeddings, semantic tokens, entities, relationships, knowledge-graph paths, semantic-layer metrics, and c Last reviewed: 2026-06-27 UTC. ### Cognitive Runtime Environment URL: https://aruntime.com/runtime-types/cognitive-runtime-environment/ Status: Emerging proposal; category: Architecture. Audience: AI architects, platform engineers, runtime implementers Search intent: Cognitive Runtime Environment definition architecture runtime Summary: A Cognitive Runtime Environment is a runtime layer that manages high-level reasoning behavior around a language model. It provides persistent memory, context paging, planning loops, task decomposition, reasoning budgets, memory validation, capability boundaries, and recovery so t Last reviewed: 2026-06-27 UTC. ### Neural Execution Engine URL: https://aruntime.com/runtime-types/neural-execution-engine/ Status: Emerging proposal; category: Architecture. Audience: AI architects, platform engineers, runtime implementers Search intent: Neural Execution Engine definition architecture runtime Summary: A Neural Execution Engine is a low-level runtime and compiler-adjacent execution layer that standardizes how neural graphs, weights, token streams, activations, memory buffers, and dispatches are mapped onto heterogeneous hardware such as CPUs, GPUs, NPUs, TPUs, DSPs, and custom Last reviewed: 2026-06-27 UTC. ## Use cases ### Enterprise Knowledge Assistant URL: https://aruntime.com/use-cases/enterprise-knowledge-assistant/ Audience: Platform architects, AI infrastructure engineers, security and governance leaders, and application developers Summary: A governed retrieval-and-generation system that answers from authorized enterprise sources, preserves tenant and data-classification boundaries, and exposes evidence for every material answer. Runtime boundary: The agentic application runtime establishes the request boundary and owns authorization-aware context assembly, model routing, citation policy, memory decisions, and evidence. The inference and serving layers still own token generation, batching, and model health. Separating these responsibilities lets the system optimize inference independently while keeping identity, data access, and answer provenance enforceable at the application boundary. Last reviewed: 2026-06-23T00:00:00Z. ### Coding Agent URL: https://aruntime.com/use-cases/coding-agent/ Audience: Application and agent developers, developer-platform teams, security engineers, and engineering managers Summary: A repository-scoped agent that can inspect code, propose patches, run bounded commands and tests, and produce reviewable evidence without receiving unconstrained workstation or credential access. Runtime boundary: The application runtime turns a coding request into a bounded execution session. It owns repository identity, branch and path permissions, sandbox configuration, tool classes, secret references, approval gates, budgets, idempotency, diffs, tests, and evidence. The model may propose commands, but the runtime decides whether a command is representable, permitted, isolated, observable, and eligible for execution. Last reviewed: 2026-06-23T00:00:00Z. ### Customer-Support Action Agent URL: https://aruntime.com/use-cases/customer-support-action-agent/ Audience: Support-platform architects, application developers, security and governance teams, and technical product leaders Summary: A support system that may read account context and propose or perform bounded actions while preserving customer identity, authorization, idempotency, human review, and customer-visible evidence. Runtime boundary: The agentic runtime separates conversational planning from account-impacting execution. It binds a verified customer and support actor to a specific case, presents only authorized tools, requires approvals at configured thresholds, validates current account state, supplies idempotency keys, and records side effects. The model server does not own these business controls; it only executes inference under the route selected by the application runtime. Last reviewed: 2026-06-23T00:00:00Z. ### Scientific or Analytical Workflow URL: https://aruntime.com/use-cases/scientific-analytical-workflow/ Audience: Researchers, technical analysts, ML systems engineers, data-platform teams, and research-governance leaders Summary: A reproducible analytical runtime that binds datasets, code, environments, intermediate artifacts, citations, validation, and result provenance into a reviewable execution record. Runtime boundary: The runtime binds each analytical claim to an admitted dataset scope, an executable plan, a controlled environment, versioned tools, intermediate artifacts, validation checks, and provenance. It separates exploratory model behavior from deterministic computation and records which statements are calculated, sourced, inferred, or proposed. This supports reproduction and review even when raw datasets or prompts cannot be broadly retained. Last reviewed: 2026-06-23T00:00:00Z. ### Local Private Assistant URL: https://aruntime.com/use-cases/local-private-assistant/ Audience: Developers evaluating local runtimes, privacy architects, desktop engineers, and technical decision-makers Summary: A device-local assistant that keeps model execution and storage on the user’s device by default, with explicit capability, update, and hosted-fallback boundaries. Runtime boundary: The local application runtime owns the boundary between device-only processing and any networked capability. It selects a compatible model and execution provider, controls local storage and memory, obtains consent for hosted fallback, mediates plugins and file access, and records whether a response was produced locally or remotely. The inference engine handles loading and token generation; the application runtime enforces the user-facing privacy contract. Last reviewed: 2026-06-23T00:00:00Z. ### Browser AI Application URL: https://aruntime.com/use-cases/browser-ai-application/ Audience: Web developers, ML systems engineers, privacy architects, and developers evaluating browser runtimes Summary: A progressively enhanced web application that executes eligible models in the browser while accounting for capability variance, model delivery, caching, privacy, offline behavior, and server fallback. Runtime boundary: The browser application runtime selects among WebGPU, WebNN where supported, WebAssembly, worker-based CPU execution, or a governed server route. It owns capability detection, model-manifest verification, origin storage, progressive download, cache eviction, privacy labeling, cancellation, and fallback. The graph or inference runtime executes the model; the application runtime handles the web lifecycle and user-visible execution contract. Last reviewed: 2026-06-23T00:00:00Z. ### Edge or Mobile System URL: https://aruntime.com/use-cases/edge-mobile-system/ Audience: Mobile and embedded engineers, ML systems engineers, edge-platform architects, and technical product leaders Summary: An on-device runtime that packages, accelerates, updates, monitors, and safely degrades models under memory, battery, thermal, connectivity, and hardware-backend constraints. Runtime boundary: The edge runtime packages the model and operators, selects CPU/GPU/NPU backends, manages memory and scheduling, and exposes health and fallback. The product application runtime adds permissions, local data policy, update rollout, telemetry minimization, and remote-service boundaries. Both layers are needed: efficient inference alone does not define update safety, privacy, or degraded behavior. Last reviewed: 2026-06-23T00:00:00Z. ## Runtime directory records ### ONNX Runtime Directory URL: https://aruntime.com/runtimes/onnx-runtime/ Canonical project URL: https://onnxruntime.ai/ Maintainer: ONNX Runtime project; availability: Open source; license: MIT. Primary category: Compiler and graph runtime; secondary categories: Inference engine. Layers: Layer 2, Layer 3. Scope: Classified primarily by graph execution responsibility; it is not a complete model-serving or agentic runtime. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://onnxruntime.ai/docs/reference/high-level-design.html; https://github.com/microsoft/onnxruntime ### vLLM Directory URL: https://aruntime.com/runtimes/vllm/ Canonical project URL: https://vllm.ai/ Maintainer: vLLM project; availability: Open source; license: Apache-2.0. Primary category: Model and LLM inference engine; secondary categories: Model server, Distributed runtime. Layers: Layer 3, Layer 4. Scope: Primary category is inference engine even when deployed behind its built-in API server. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.vllm.ai/; https://github.com/vllm-project/vllm ### SGLang Directory URL: https://aruntime.com/runtimes/sglang/ Canonical project URL: https://docs.sglang.ai/ Maintainer: SGLang project; availability: Open source; license: Apache-2.0. Primary category: Model and LLM inference engine; secondary categories: Structured generation runtime, Model server. Layers: Layer 3, Layer 4. Scope: Structured generation is not equivalent to a governed agentic application runtime. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.sglang.ai/; https://github.com/sgl-project/sglang ### NVIDIA Triton Inference Server Directory URL: https://aruntime.com/runtimes/nvidia-triton-inference-server/ Canonical project URL: https://developer.nvidia.com/triton-inference-server Maintainer: NVIDIA; availability: Open source; license: BSD-3-Clause. Primary category: Model serving runtime; secondary categories: Serving platform. Layers: Layer 4. Scope: A model server does not automatically provide agent tool authorization, durable application state, or business workflow. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/architecture.html; https://github.com/triton-inference-server/server ### NVIDIA TensorRT Directory URL: https://aruntime.com/runtimes/nvidia-tensorrt/ Canonical project URL: https://developer.nvidia.com/tensorrt Maintainer: NVIDIA; availability: Commercial SDK with public components; license: NVIDIA software license; public components vary. Primary category: Compiler and graph runtime; secondary categories: Inference engine. Layers: Layer 2, Layer 3. Scope: Classified across compilation and inference because optimization produces an executable engine used by the runtime. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.nvidia.com/deeplearning/tensorrt/latest/; https://github.com/NVIDIA/TensorRT ### TensorRT-LLM Directory URL: https://aruntime.com/runtimes/tensorrt-llm/ Canonical project URL: https://nvidia.github.io/TensorRT-LLM/ Maintainer: NVIDIA; availability: Open source; license: Apache-2.0. Primary category: Model and LLM inference engine; secondary categories: Compiler and graph runtime, Distributed runtime. Layers: Layer 2, Layer 3, Layer 4. Scope: The serving layer depends on how the engine is packaged and deployed. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://nvidia.github.io/TensorRT-LLM/; https://github.com/NVIDIA/TensorRT-LLM ### Apache TVM Directory URL: https://aruntime.com/runtimes/apache-tvm/ Canonical project URL: https://tvm.apache.org/ Maintainer: Apache Software Foundation; availability: Open source; license: Apache-2.0. Primary category: Compiler and graph runtime; secondary categories: Kernel generation and scheduling. Layers: Layer 1, Layer 2. Scope: TVM spans compiler and low-level runtime responsibilities but is not a model-serving platform by itself. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://tvm.apache.org/docs/; https://github.com/apache/tvm ### StableHLO Directory URL: https://aruntime.com/runtimes/stablehlo/ Canonical project URL: https://openxla.org/stablehlo Maintainer: OpenXLA project; availability: Open source specification and implementation; license: Apache-2.0. Primary category: Intermediate representation; secondary categories: Compiler ecosystem component. Layers: Layer 2. Scope: Included as a boundary example: an IR is a critical runtime-stack component but not a complete runtime product. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://openxla.org/stablehlo/spec; https://github.com/openxla/stablehlo ### ExecuTorch Directory URL: https://aruntime.com/runtimes/executorch/ Canonical project URL: https://pytorch.org/executorch/ Maintainer: PyTorch project; availability: Open source; license: BSD-3-Clause. Primary category: Edge and mobile inference runtime; secondary categories: Compiler and graph runtime. Layers: Layer 2, Layer 3. Scope: On-device packaging, delegate selection, memory limits, and update policy are part of the deployment architecture. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.pytorch.org/executorch/stable/intro-overview.html; https://github.com/pytorch/executorch ### LiteRT Directory URL: https://aruntime.com/runtimes/litert/ Canonical project URL: https://ai.google.dev/edge/litert Maintainer: Google AI Edge; availability: Open source components; license: Apache-2.0 for reviewed repository components. Primary category: Edge and mobile inference runtime; secondary categories: Inference engine. Layers: Layer 3. Scope: The name and migration path from TensorFlow Lite should be verified against the exact release being adopted. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://ai.google.dev/edge/litert; https://github.com/google-ai-edge/LiteRT ### WebNN Directory URL: https://aruntime.com/runtimes/webnn/ Canonical project URL: https://www.w3.org/TR/webnn/ Maintainer: W3C Web Machine Learning Community Group and Working Group participants; availability: Web specification; license: W3C document and software terms. Primary category: Browser inference API; secondary categories: Graph runtime interface. Layers: Layer 2, Layer 3. Scope: A specification is not itself a complete product runtime. Browser implementation status and supported operations must be tested. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://www.w3.org/TR/webnn/; https://github.com/webmachinelearning/webnn ### WebGPU Directory URL: https://aruntime.com/runtimes/webgpu/ Canonical project URL: https://www.w3.org/TR/webgpu/ Maintainer: W3C GPU for the Web Working Group; availability: Web specification; license: W3C document and software terms. Primary category: Browser compute API; secondary categories: Hardware abstraction used by inference libraries. Layers: Layer 1, Layer 2, Layer 3. Scope: WebGPU is a general compute substrate, not an inference engine or model server. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://www.w3.org/TR/webgpu/; https://github.com/gpuweb/gpuweb ### OpenVINO Directory URL: https://aruntime.com/runtimes/openvino/ Canonical project URL: https://docs.openvino.ai/ Maintainer: Intel; availability: Open source; license: Apache-2.0. Primary category: Compiler and graph runtime; secondary categories: Inference engine. Layers: Layer 2, Layer 3. Scope: Classified across graph optimization and inference; surrounding serving and governance are separate layers. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.openvino.ai/; https://github.com/openvinotoolkit/openvino ### KServe Directory URL: https://aruntime.com/runtimes/kserve/ Canonical project URL: https://kserve.github.io/website/ Maintainer: KServe project; availability: Open source; license: Apache-2.0. Primary category: Serving platform; secondary categories: Model serving runtime, Kubernetes operator. Layers: Layer 4. Scope: The platform orchestrates serving runtimes; it does not make all backends equivalent or own agentic tool policy. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://kserve.github.io/website/; https://github.com/kserve/kserve ### Ray Serve Directory URL: https://aruntime.com/runtimes/ray-serve/ Canonical project URL: https://docs.ray.io/en/latest/serve/ Maintainer: Ray project; availability: Open source; license: Apache-2.0. Primary category: Model serving runtime; secondary categories: Distributed application runtime. Layers: Layer 4. Scope: A general distributed application graph may include model calls, but agentic governance remains an application responsibility. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.ray.io/en/latest/serve/; https://github.com/ray-project/ray ### BentoML Directory URL: https://aruntime.com/runtimes/bentoml/ Canonical project URL: https://docs.bentoml.com/ Maintainer: BentoML project; availability: Open source with commercial ecosystem; license: Apache-2.0 for BentoML repository. Primary category: Model serving framework; secondary categories: Packaging and deployment tooling. Layers: Layer 4. Scope: Packaging and serving do not imply a universal inference backend or agentic governance layer. Last verified: 2026-06-23T00:00:00Z. Authoritative sources: https://docs.bentoml.com/; https://github.com/bentoml/BentoML ## Glossary ### Admission control Anchor: https://aruntime.com/glossary/#admission-control Category: Serving and distributed execution; layer: Serving and distributed execution. A decision at the service boundary to accept, queue, shed, defer, or reject work based on capacity, authority, priority, or policy. Related: Dynamic batching, Backpressure, Goodput, Tensor parallelism. ### Agent framework Anchor: https://aruntime.com/glossary/#agent-framework Category: Agentic and application runtime; layer: Agentic and application runtime. A programming framework that supplies abstractions for model calls, tools, state, planning, or multi-step flows; it may participate in but does not by itself guarantee production runtime controls. Related: Agentic runtime, Tool broker, Tool contract, Side effect. ### Agentic Execution Environment Anchor: https://aruntime.com/glossary/#agentic-execution-environment Category: Agentic; layer: L5. A runtime boundary for safely running autonomous or semi-autonomous agent work with sandboxing, lifecycle control, tool authorization, secret brokering, audit, and recovery. Related: agentic runtime, sandbox, policy hook. ### Agentic runtime Anchor: https://aruntime.com/glossary/#agentic-runtime Category: Agentic and application runtime; layer: Agentic and application runtime. An application-runtime layer that turns probabilistic model behavior into controlled work by governing request boundaries, identity, context, tools, memory, policy, approvals, evaluation, evidence, and recovery. Related: Agent framework, Tool broker, Tool contract, Side effect. ### AI runtime Anchor: https://aruntime.com/glossary/#ai-runtime Category: Runtime fundamentals; layer: Runtime fundamentals. The execution environment that turns model artifacts or model requests into operational behavior; depending on the layer, it may compile graphs, schedule hardware, execute inference, serve models, coordinate distributed workloads, or govern context, tools, memory, policy, and traces. Related: Inference runtime, Inference engine, Model server, Serving platform. ### AI-Ready Web Anchor: https://aruntime.com/glossary/#ai-ready-web Category: Governance and discovery; layer: Cross-layer. A public-content readiness pattern where visible pages, metadata, discovery files, correction paths, and support boundaries agree. Related: Answer Engine Optimization, Generative Engine Optimization, SEO, llms.txt. ### Answer Engine Optimization Anchor: https://aruntime.com/glossary/#answer-engine-optimization Category: Discovery and publishing; layer: Product and workflow layer. AEO is the practice of publishing direct, citable, source-backed answers that answer engines can extract and route back to the canonical page. Related: Generative Engine Optimization, AI-Ready Web, Canonical URL. ### AOT compilation Anchor: https://aruntime.com/glossary/#aot-compilation Category: Compiler and graph; layer: Compiler and graph. Compilation performed before deployment or execution to create target-ready artifacts with controlled versions and compatibility assumptions. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Application runtime Anchor: https://aruntime.com/glossary/#application-runtime Category: Runtime fundamentals; layer: Runtime fundamentals. An execution environment that coordinates application-level state, policies, dependencies, and work around lower-level model execution. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Approval gate Anchor: https://aruntime.com/glossary/#approval-gate Category: Agentic and application runtime; layer: Agentic and application runtime. A runtime checkpoint that blocks an action until an authorized human or system supplies a time-bounded decision. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Autonomous Inference Runtime Anchor: https://aruntime.com/glossary/#autonomous-inference-runtime Category: Edge; layer: L0-L4. A local or edge-oriented runtime that executes trained models near the device, sensor, user, robot, vehicle, or industrial process that needs the result. Related: edge inference, inference engine, model server. ### Autoscaling Anchor: https://aruntime.com/glossary/#autoscaling Category: Serving and distributed execution; layer: Serving and distributed execution. Adjustment of runtime capacity using demand, queue, utilization, latency, memory, or custom workload signals. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Backpressure Anchor: https://aruntime.com/glossary/#backpressure Category: Serving and distributed execution; layer: Serving and distributed execution. A mechanism that limits upstream work when downstream capacity, queues, memory, or dependencies are saturated. Related: Dynamic batching, Admission control, Goodput, Tensor parallelism. ### Benchmark fixture Anchor: https://aruntime.com/glossary/#benchmark-fixture Category: Observability and evaluation; layer: Observability and evaluation. A versioned set of models, inputs, expected properties, environment assumptions, and measurement instructions used for repeatable evaluation. Related: Runtime trace, Span, Trace correlation, Replay. ### Bitemporal memory Anchor: https://aruntime.com/glossary/#bitemporal-memory Category: Memory; layer: L4-L5. A memory or graph model that records both when a fact was valid in the world and when the system recorded or learned it. Related: enterprise memory graph, provenance. ### Browser runtime Anchor: https://aruntime.com/glossary/#browser-runtime Category: Edge, browser, and local; layer: Edge, browser, and local. A model-execution environment operating within browser security and capability boundaries, commonly using JavaScript, WebAssembly, WebGPU, or platform APIs. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### Business outcome Anchor: https://aruntime.com/glossary/#business-outcome Category: Observability and evaluation; layer: Observability and evaluation. A product- or domain-level result, such as a resolved case or accepted change, that should be measured separately from model output quality. Related: Runtime trace, Span, Trace correlation, Replay. ### Canary rollout Anchor: https://aruntime.com/glossary/#canary-rollout Category: Serving and distributed execution; layer: Serving and distributed execution. A deployment strategy that sends a limited, measured share of traffic to a new version before broader promotion. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Chunked prefill Anchor: https://aruntime.com/glossary/#chunked-prefill Category: Inference and memory; layer: Inference and memory. Execution of long prompt prefill in bounded chunks so prefill work can be interleaved with decode or other requests. Related: Prefill, Decode, KV cache, PagedAttention. ### Code generation Anchor: https://aruntime.com/glossary/#code-generation Category: Compiler and graph; layer: Compiler and graph. The production of executable machine code, kernels, bytecode, or target artifacts from an intermediate representation. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Cognitive Runtime Environment Anchor: https://aruntime.com/glossary/#cognitive-runtime-environment Category: Agentic; layer: L5. A runtime layer that manages high-level reasoning behavior around a language model, including persistent memory, context paging, planning loops, reasoning budgets, and recovery. Related: context window, memory retrieval, planning loop. ### Cold start Anchor: https://aruntime.com/glossary/#cold-start Category: Serving and distributed execution; layer: Serving and distributed execution. Latency and resource work incurred when an execution environment, model, cache, or worker is not already initialized. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Collective communication Anchor: https://aruntime.com/glossary/#collective-communication Category: Serving and distributed execution; layer: Serving and distributed execution. A coordinated multi-participant operation such as all-reduce, all-gather, reduce-scatter, or broadcast used in distributed execution. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Compensation Anchor: https://aruntime.com/glossary/#compensation Category: Agentic and application runtime; layer: Agentic and application runtime. A forward action intended to counteract or mitigate a completed side effect when transactional rollback is unavailable. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Compiler runtime Anchor: https://aruntime.com/glossary/#compiler-runtime Category: Runtime fundamentals; layer: Runtime fundamentals. The runtime portion of a compiler-backed system that loads compiled artifacts, binds inputs, allocates memory, dispatches generated code, and handles supported fallbacks. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Confused deputy Anchor: https://aruntime.com/glossary/#confused-deputy Category: Security, policy, and evidence; layer: Security, policy, and evidence. A security failure in which a component with authority is induced to misuse that authority on behalf of a less-privileged requester. Related: Prompt injection, Policy decision point, Policy enforcement point, Data classification. ### Context assembly Anchor: https://aruntime.com/glossary/#context-assembly Category: Agentic and application runtime; layer: Agentic and application runtime. Selection, authorization, transformation, ordering, and attribution of instructions, retrieved data, state, and other inputs supplied to execution. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Context length Anchor: https://aruntime.com/glossary/#context-length Category: Inference and memory; layer: Inference and memory. The number of tokens or other units a model accepts within its configured input and generated sequence window. Related: Prefill, Decode, KV cache, PagedAttention. ### Continuous batching Anchor: https://aruntime.com/glossary/#continuous-batching Category: Inference and memory; layer: Inference and memory. A scheduling method that admits and removes sequences while a batch is running instead of waiting for all batch members to finish. Related: Prefill, Decode, KV cache, PagedAttention. ### Cost per successful task Anchor: https://aruntime.com/glossary/#cost-per-successful-task Category: Observability and evaluation; layer: Observability and evaluation. Total scoped runtime cost divided by tasks that satisfy defined quality, policy, and completion criteria. Related: Runtime trace, Span, Trace correlation, Replay. ### Data classification Anchor: https://aruntime.com/glossary/#data-classification Category: Security, policy, and evidence; layer: Security, policy, and evidence. Assignment of handling categories to data so access, routing, retention, logging, and egress rules can be enforced. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Decode Anchor: https://aruntime.com/glossary/#decode Category: Inference and memory; layer: Inference and memory. The iterative LLM inference phase that generates subsequent tokens while reading and extending the key-value cache. Related: Prefill, KV cache, PagedAttention, Prefix caching. ### Delegate Anchor: https://aruntime.com/glossary/#delegate Category: Compiler and graph; layer: Compiler and graph. A component that receives supported operations from a host runtime and executes them on a specialized backend or accelerator. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Disaggregated prefill and decode Anchor: https://aruntime.com/glossary/#disaggregated-prefill-and-decode Category: Serving and distributed execution; layer: Serving and distributed execution. An architecture in which prompt processing and token decoding run on separately scaled workers with model state transferred or remotely accessed. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Dream cycle Anchor: https://aruntime.com/glossary/#dream-cycle Category: Memory; layer: L5. An offline or background consolidation process that deduplicates, compresses, connects, archives, or retires memories. Related: memory consolidation, cognitive runtime environment. ### Dual attribution Anchor: https://aruntime.com/glossary/#dual-attribution Category: Governance; layer: L5-L6. A logging and identity pattern that records both the machine actor and the authenticated human authority behind a workflow. Related: evidence schema, machine intelligence runtime. ### Dynamic batching Anchor: https://aruntime.com/glossary/#dynamic-batching Category: Serving and distributed execution; layer: Serving and distributed execution. Server-side combination of requests into batches according to queue, delay, shape, priority, and model constraints. Related: Admission control, Backpressure, Goodput, Tensor parallelism. ### Dynamic shape Anchor: https://aruntime.com/glossary/#dynamic-shape Category: Compiler and graph; layer: Compiler and graph. A tensor dimension whose value may vary between executions rather than being fixed in the compiled artifact. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Edge runtime Anchor: https://aruntime.com/glossary/#edge-runtime Category: Edge, browser, and local; layer: Edge, browser, and local. A runtime deployed near data sources or users under constrained connectivity, resource, latency, update, and physical-security conditions. Related: Local runtime, On-device runtime, TinyML, Browser runtime. ### Evaluation result Anchor: https://aruntime.com/glossary/#evaluation-result Category: Observability and evaluation; layer: Observability and evaluation. A versioned record of a test case, configuration, grader or rule, score or decision, threshold, evidence, and pass/fail interpretation. Related: Runtime trace, Span, Trace correlation, Replay. ### Evidence record Anchor: https://aruntime.com/glossary/#evidence-record Category: Security, policy, and evidence; layer: Security, policy, and evidence. A minimized, correlated record of request identity, versions, context references, model route, tool actions, policy and approval decisions, side effects, failures, recovery, evaluation, redaction, and retention. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Execution boundary Anchor: https://aruntime.com/glossary/#execution-boundary Category: Runtime fundamentals; layer: Runtime fundamentals. The point at which ownership of inputs, authority, state, failures, telemetry, or side effects transfers between runtime components. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Execution provider Anchor: https://aruntime.com/glossary/#execution-provider Category: Compiler; layer: L1-L3. A backend that advertises which neural graph operations it can execute on a target device or library. Related: ONNX Runtime, backend. ### Expert parallelism Anchor: https://aruntime.com/glossary/#expert-parallelism Category: Serving and distributed execution; layer: Serving and distributed execution. Distribution and routing of mixture-of-experts components across devices or hosts. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Fallback Anchor: https://aruntime.com/glossary/#fallback Category: Compiler and graph; layer: Compiler and graph. Execution on a less-specialized path when an operation, shape, device, or feature is unsupported by the preferred backend. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Framework Anchor: https://aruntime.com/glossary/#framework Category: Runtime fundamentals; layer: Runtime fundamentals. A programming environment that supplies model construction, training, export, execution, or application abstractions; a framework may embed a runtime but is not automatically one category. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Generative Engine Optimization Anchor: https://aruntime.com/glossary/#generative-engine-optimization Category: Discovery and publishing; layer: Product and workflow layer. GEO is the practice of publishing retrievable, grounded, source-resilient content that generative systems can synthesize and cite without losing context or boundaries. Related: Answer Engine Optimization, AI-Ready Web, Provenance. ### Goodput Anchor: https://aruntime.com/glossary/#goodput Category: Serving and distributed execution; layer: Serving and distributed execution. The rate of requests or tokens that meet defined quality and service-level requirements, rather than raw completed throughput alone. Related: Dynamic batching, Admission control, Backpressure, Tensor parallelism. ### Graph capture Anchor: https://aruntime.com/glossary/#graph-capture Category: Compiler and graph; layer: Compiler and graph. The process of converting observed or declared model operations into a graph or program representation suitable for analysis and optimization. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Graph partitioning Anchor: https://aruntime.com/glossary/#graph-partitioning Category: Compiler and graph; layer: Compiler and graph. The division of a graph into subgraphs assigned to different devices, backends, execution providers, or fallback paths. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Graph rewrite Anchor: https://aruntime.com/glossary/#graph-rewrite Category: Compiler and graph; layer: Compiler and graph. A semantics-preserving transformation that replaces part of a computation graph with an equivalent form. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Graph runtime Anchor: https://aruntime.com/glossary/#graph-runtime Category: Runtime fundamentals; layer: Runtime fundamentals. An execution environment that evaluates a model or computation graph, often after graph rewriting, partitioning, scheduling, and memory planning. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Human review Anchor: https://aruntime.com/glossary/#human-review Category: Agentic and application runtime; layer: Agentic and application runtime. A defined process in which an authorized person evaluates a proposed action, output, exception, or recovery path before execution or release. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Idempotency key Anchor: https://aruntime.com/glossary/#idempotency-key Category: Agentic and application runtime; layer: Agentic and application runtime. A stable identifier that lets a system recognize repeated attempts for the same logical operation and avoid unintended duplicate effects. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Inference engine Anchor: https://aruntime.com/glossary/#inference-engine Category: Runtime fundamentals; layer: Runtime fundamentals. A system optimized to load model weights and perform inference, including device execution, memory management, batching, streaming, and model-specific generation mechanics. Related: AI runtime, Inference runtime, Model server, Serving platform. ### Inference runtime Anchor: https://aruntime.com/glossary/#inference-runtime Category: Runtime fundamentals; layer: Runtime fundamentals. A runtime whose primary responsibility is executing trained model artifacts to produce predictions, embeddings, tokens, or other model outputs. Related: AI runtime, Inference engine, Model server, Serving platform. ### Intermediate representation Anchor: https://aruntime.com/glossary/#intermediate-representation Category: Compiler and graph; layer: Compiler and graph. A compiler data model used between source or framework graphs and target code so transformations, analysis, partitioning, and lowering can be performed systematically. Related: ONNX, MLIR, StableHLO, TOSA. ### ITL Anchor: https://aruntime.com/glossary/#itl Category: Inference and memory; layer: Inference and memory. Inter-token latency: the observed latency between individual streamed output tokens, often analyzed as a distribution rather than a single mean. Related: Prefill, Decode, KV cache, PagedAttention. ### JIT compilation Anchor: https://aruntime.com/glossary/#jit-compilation Category: Compiler and graph; layer: Compiler and graph. Compilation performed during or near execution, commonly specialized to observed shapes, values, devices, or runtime conditions. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### KV cache Anchor: https://aruntime.com/glossary/#kv-cache Category: Inference and memory; layer: Inference and memory. Stored attention keys and values from prior tokens that avoid recomputing the full sequence during autoregressive decoding. Related: Prefill, Decode, PagedAttention, Prefix caching. ### KV cache orchestration Anchor: https://aruntime.com/glossary/#kv-cache-orchestration Category: Inference; layer: L3-L4. Runtime management of transformer key-value cache allocation, reuse, paging, sharing, and offload. Related: KV cache, LLM inference. ### Library Anchor: https://aruntime.com/glossary/#library Category: Runtime fundamentals; layer: Runtime fundamentals. A reusable collection of functions or components invoked by another program; it does not necessarily own process, request, scheduling, or lifecycle boundaries. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Local runtime Anchor: https://aruntime.com/glossary/#local-runtime Category: Edge, browser, and local; layer: Edge, browser, and local. A runtime executing on the user or operator device rather than a remote hosted service. Related: Edge runtime, On-device runtime, TinyML, Browser runtime. ### Long-term memory Anchor: https://aruntime.com/glossary/#long-term-memory Category: Agentic and application runtime; layer: Agentic and application runtime. Durable state made available across tasks or sessions under explicit provenance, authorization, update, deletion, and retention policies. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Lowering Anchor: https://aruntime.com/glossary/#lowering Category: Compiler and graph; layer: Compiler and graph. The transformation of a high-level operation or IR into a more target-specific representation or implementation. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Machine Intelligence Runtime Anchor: https://aruntime.com/glossary/#machine-intelligence-runtime Category: Enterprise; layer: L3-L6. An enterprise execution-time control layer that coordinates model-backed work across identity, policy, memory, tools, model routing, telemetry, audit, budgets, and governed state changes. Related: runtime contract, evidence schema, agentic runtime. ### Memory planning Anchor: https://aruntime.com/glossary/#memory-planning Category: Compiler and graph; layer: Compiler and graph. The compile-time or runtime assignment and reuse of buffers based on tensor lifetimes, layouts, shapes, and device constraints. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Memory pressure Anchor: https://aruntime.com/glossary/#memory-pressure Category: Inference and memory; layer: Inference and memory. A condition in which model weights, activations, caches, buffers, or other allocations approach available memory and force eviction, offload, throttling, or failure. Related: Prefill, Decode, KV cache, PagedAttention. ### Mixed precision Anchor: https://aruntime.com/glossary/#mixed-precision Category: Inference and memory; layer: Inference and memory. Use of multiple numeric precisions within one model or execution path to balance accuracy, speed, and resource use. Related: Prefill, Decode, KV cache, PagedAttention. ### MLIR Anchor: https://aruntime.com/glossary/#mlir Category: Compiler and graph; layer: Compiler and graph. A compiler infrastructure for defining and transforming multiple intermediate-representation dialects across abstraction levels. Related: Intermediate representation, ONNX, StableHLO, TOSA. ### Model artifact integrity Anchor: https://aruntime.com/glossary/#model-artifact-integrity Category: Security, policy, and evidence; layer: Security, policy, and evidence. Verification that model files and associated assets match approved hashes, signatures, provenance, versions, and supply-chain policy. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Model lifecycle Anchor: https://aruntime.com/glossary/#model-lifecycle Category: Runtime fundamentals; layer: Runtime fundamentals. The sequence from model creation and export through optimization, packaging, deployment, loading, serving, version transition, and retirement. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Model packaging Anchor: https://aruntime.com/glossary/#model-packaging Category: Edge, browser, and local; layer: Edge, browser, and local. Assembly of weights, tokenizer or preprocessing assets, runtime configuration, licenses, integrity data, and compatibility metadata for deployment. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### Model repository Anchor: https://aruntime.com/glossary/#model-repository Category: Serving and distributed execution; layer: Serving and distributed execution. A versioned location and layout from which a serving system discovers and loads model artifacts and configuration. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Model server Anchor: https://aruntime.com/glossary/#model-server Category: Runtime fundamentals; layer: Runtime fundamentals. A networked service that exposes one or more models through APIs and adds loading, health, scheduling, batching, version, and repository behavior around an inference engine. Related: AI runtime, Inference runtime, Inference engine, Serving platform. ### Multi-model serving Anchor: https://aruntime.com/glossary/#multi-model-serving Category: Serving and distributed execution; layer: Serving and distributed execution. Operation of several models or versions within a shared serving fleet, often requiring placement, isolation, loading, eviction, and routing policies. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Neural Execution Engine Anchor: https://aruntime.com/glossary/#neural-execution-engine Category: Compiler; layer: L1-L3. A low-level runtime and compiler-adjacent execution layer that maps neural graphs, weights, activations, and token steps onto heterogeneous hardware. Related: execution provider, intermediate representation, kernel. ### Neural Runtime Abstraction Layer Anchor: https://aruntime.com/glossary/#neural-runtime-abstraction-layer Category: Compiler; layer: L1-L3. An abstraction layer that hides hardware-specific neural execution details behind a common runtime interface. Related: neural execution engine, hardware abstraction layer. ### No-op safety Anchor: https://aruntime.com/glossary/#no-op-safety Category: Governance and discovery; layer: Agentic and application runtime. A safety posture where unsupported, unauthorized, or unsafe requests stop cleanly instead of guessing, escalating authority, or executing a side effect. Related: Policy hook, Support boundary, Human approval. ### Offline inference Anchor: https://aruntime.com/glossary/#offline-inference Category: Edge, browser, and local; layer: Edge, browser, and local. Model execution that does not require an active network connection after necessary artifacts and configuration are available locally. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### On-device runtime Anchor: https://aruntime.com/glossary/#on-device-runtime Category: Edge, browser, and local; layer: Edge, browser, and local. A runtime that performs model execution within a phone, desktop, embedded device, appliance, or other endpoint. Related: Local runtime, Edge runtime, TinyML, Browser runtime. ### ONNX Anchor: https://aruntime.com/glossary/#onnx Category: Compiler and graph; layer: Compiler and graph. An open format for representing machine-learning model graphs and operators so models can be exchanged among compatible tools and runtimes. Related: Intermediate representation, MLIR, StableHLO, TOSA. ### Opaque secret broker Anchor: https://aruntime.com/glossary/#opaque-secret-broker Category: Security; layer: L5. A runtime service that gives an agent a handle to a credential rather than placing the raw secret inside the agent environment. Related: secret broker, agentic execution environment. ### Operator fusion Anchor: https://aruntime.com/glossary/#operator-fusion Category: Compiler and graph; layer: Compiler and graph. The combination of adjacent graph operations into a larger operation to reduce intermediate memory traffic or dispatch overhead. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Overthinking intervention Anchor: https://aruntime.com/glossary/#overthinking-intervention Category: Agentic; layer: L5. A runtime action that terminates, redirects, verifies, or escalates reasoning when traces show oscillation, repeated reconsideration, or rising cost without progress. Related: reasoning budget, cognitive runtime environment. ### PagedAttention Anchor: https://aruntime.com/glossary/#pagedattention Category: Inference and memory; layer: Inference and memory. A KV-cache management technique that maps logical cache blocks to non-contiguous physical memory blocks to improve allocation and sharing. Related: Prefill, Decode, KV cache, Prefix caching. ### Pipeline parallelism Anchor: https://aruntime.com/glossary/#pipeline-parallelism Category: Serving and distributed execution; layer: Serving and distributed execution. Distribution of sequential model stages across devices or workers, with microbatches moving through the pipeline. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Policy decision point Anchor: https://aruntime.com/glossary/#policy-decision-point Category: Security, policy, and evidence; layer: Security, policy, and evidence. A component that evaluates policy inputs and returns an explicit decision and reason codes. Related: Prompt injection, Confused deputy, Policy enforcement point, Data classification. ### Policy enforcement point Anchor: https://aruntime.com/glossary/#policy-enforcement-point Category: Security, policy, and evidence; layer: Security, policy, and evidence. A component that applies a policy decision at the boundary where access, routing, data release, or action occurs. Related: Prompt injection, Confused deputy, Policy decision point, Data classification. ### Policy hook Anchor: https://aruntime.com/glossary/#policy-hook Category: Security; layer: L5. A deterministic interception point that evaluates a proposed action before or after execution. Related: policy engine, tool authorization. ### Post-training quantization Anchor: https://aruntime.com/glossary/#post-training-quantization Category: Inference and memory; layer: Inference and memory. Quantization applied after model training using calibration, statistics, or weight-only conversion rather than retraining the full model. Related: Prefill, Decode, KV cache, PagedAttention. ### Prefill Anchor: https://aruntime.com/glossary/#prefill Category: Inference and memory; layer: Inference and memory. The LLM inference phase that processes the input sequence, produces the first-token state, and populates key-value cache entries for subsequent decoding. Related: Decode, KV cache, PagedAttention, Prefix caching. ### Prefix caching Anchor: https://aruntime.com/glossary/#prefix-caching Category: Inference and memory; layer: Inference and memory. Reuse of previously computed model state for an identical or compatible input prefix. Related: Prefill, Decode, KV cache, PagedAttention. ### Progressive fallback Anchor: https://aruntime.com/glossary/#progressive-fallback Category: Edge, browser, and local; layer: Edge, browser, and local. A design that selects the best supported local path and falls back to another local or hosted path while preserving explicit privacy and behavior rules. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### Prompt injection Anchor: https://aruntime.com/glossary/#prompt-injection Category: Security, policy, and evidence; layer: Security, policy, and evidence. Untrusted content crafted to alter model behavior or override intended instructions, especially when the model can access tools, data, or privileged actions. Related: Confused deputy, Policy decision point, Policy enforcement point, Data classification. ### Provenance Anchor: https://aruntime.com/glossary/#provenance Category: Security, policy, and evidence; layer: Security, policy, and evidence. Information describing the origin, custody, transformation, version, and authority of data, models, instructions, outputs, or evidence. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Quality regression Anchor: https://aruntime.com/glossary/#quality-regression Category: Observability and evaluation; layer: Observability and evaluation. A measurable decline in task quality, validity, safety, or outcome metrics relative to an approved baseline. Related: Runtime trace, Span, Trace correlation, Replay. ### Quantization Anchor: https://aruntime.com/glossary/#quantization Category: Inference and memory; layer: Inference and memory. Representation and execution of model values at reduced precision or with compressed numeric formats to lower memory, bandwidth, or compute cost. Related: Prefill, Decode, KV cache, PagedAttention. ### Reasoning budget Anchor: https://aruntime.com/glossary/#reasoning-budget Category: Agentic; layer: L5. A bound on tokens, time, steps, branches, or tool calls that a reasoning loop may consume before termination or escalation. Related: cognitive runtime environment, token budget. ### Redaction Anchor: https://aruntime.com/glossary/#redaction Category: Security, policy, and evidence; layer: Security, policy, and evidence. Removal, masking, tokenization, or transformation of sensitive data before storage, display, telemetry, or downstream transfer. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Replay Anchor: https://aruntime.com/glossary/#replay Category: Observability and evaluation; layer: Observability and evaluation. Re-execution or reconstruction of a prior request using preserved versions, inputs or references, decisions, and controlled substitutes to diagnose or evaluate behavior. Related: Runtime trace, Span, Trace correlation, Tail latency. ### Request lifecycle Anchor: https://aruntime.com/glossary/#request-lifecycle Category: Runtime fundamentals; layer: Runtime fundamentals. The sequence from request admission through identity, context, routing, execution, validation, policy, response, evidence, evaluation, and recovery. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Retention policy Anchor: https://aruntime.com/glossary/#retention-policy Category: Security, policy, and evidence; layer: Security, policy, and evidence. Rules governing how long specific data and evidence categories are kept, where they may be stored, and how they are deleted or archived. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Route inventory Anchor: https://aruntime.com/glossary/#route-inventory Category: Discovery and publishing; layer: Product and workflow layer. A reviewed list of public routes with titles, purposes, review dates, indexability, and support boundaries. Related: Sitemap, Content catalog, llms.txt. ### Runtime request contract Anchor: https://aruntime.com/glossary/#runtime-request-contract Category: Agentic and application runtime; layer: Agentic and application runtime. A versioned envelope that states actor, tenant, task, risk, permissions, context, route, tools, memory, budget, approval, output, trace, classification, retention, idempotency, and deadline requirements. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Runtime trace Anchor: https://aruntime.com/glossary/#runtime-trace Category: Observability and evaluation; layer: Observability and evaluation. A correlated representation of work across runtime components and layers, using spans, events, links, attributes, and status while respecting data-minimization rules. Related: Span, Trace correlation, Replay, Tail latency. ### Sandbox Anchor: https://aruntime.com/glossary/#sandbox Category: Security, policy, and evidence; layer: Security, policy, and evidence. A constrained execution environment that limits code, filesystem, network, process, credential, and device access. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Search Engine Optimization Anchor: https://aruntime.com/glossary/#search-engine-optimization Category: Discovery and publishing; layer: Product and workflow layer. SEO is the practice of making public pages crawlable, understandable, canonical, internally linked, performant, accessible, and useful for search users. Related: Canonical URL, Sitemap, Structured data. ### Semantic compiler Anchor: https://aruntime.com/glossary/#semantic-compiler Category: Semantic; layer: L4-L5. A runtime component that transforms natural-language intent into governed SQL, API calls, code, or structured execution plans. Related: semantic runtime engine, semantic layer. ### Semantic control plane Anchor: https://aruntime.com/glossary/#semantic-control-plane Category: Security; layer: L4-L5. A policy layer that binds access and execution controls to semantic concepts rather than only physical tables, files, or API routes. Related: policy engine, semantic runtime engine. ### Semantic Runtime Engine Anchor: https://aruntime.com/glossary/#semantic-runtime-engine Category: Semantic; layer: L4-L5. A runtime layer that manages meaning as an operational data type, including embeddings, semantic tokens, graph relationships, semantic-layer metrics, and compiled intent. Related: vector database, knowledge graph, semantic layer. ### Semantic token Anchor: https://aruntime.com/glossary/#semantic-token Category: Semantic; layer: L4-L5. A discrete representation of meaning used to stabilize or categorize otherwise continuous semantic embeddings or concepts. Related: embedding, ontology. ### Serving platform Anchor: https://aruntime.com/glossary/#serving-platform Category: Runtime fundamentals; layer: Runtime fundamentals. An operational layer that deploys and manages model servers across environments, including rollout, autoscaling, traffic, multi-model placement, policy, and fleet health. Related: AI runtime, Inference runtime, Inference engine, Model server. ### Shape guard Anchor: https://aruntime.com/glossary/#shape-guard Category: Compiler and graph; layer: Compiler and graph. A condition that determines whether a compiled specialization is valid for a given input shape or value constraint. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### Side effect Anchor: https://aruntime.com/glossary/#side-effect Category: Agentic and application runtime; layer: Agentic and application runtime. A change outside the model’s transient computation, such as writing a record, sending a message, executing code, moving money, or changing infrastructure. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ### Side-effect classification Anchor: https://aruntime.com/glossary/#side-effect-classification Category: Agentic; layer: L5. A method for classifying tools or actions by whether they read, write, send, spend, delete, actuate, or otherwise change external state. Related: tool contract, agentic execution environment. ### Span Anchor: https://aruntime.com/glossary/#span Category: Observability and evaluation; layer: Observability and evaluation. A timed unit of work within a trace, identified by trace and span identifiers and annotated with bounded attributes, events, status, and relationships. Related: Runtime trace, Trace correlation, Replay, Tail latency. ### Speculative decoding Anchor: https://aruntime.com/glossary/#speculative-decoding Category: Inference and memory; layer: Inference and memory. Generation that uses one or more draft mechanisms to propose tokens that a target model verifies, with the goal of reducing decode latency. Related: Prefill, Decode, KV cache, PagedAttention. ### StableHLO Anchor: https://aruntime.com/glossary/#stablehlo Category: Compiler and graph; layer: Compiler and graph. A portable operation set and serialization format in the OpenXLA ecosystem intended to preserve compatibility for machine-learning compiler workflows. Related: Intermediate representation, ONNX, MLIR, TOSA. ### Structured generation Anchor: https://aruntime.com/glossary/#structured-generation Category: Inference and memory; layer: Inference and memory. Generation constrained or validated against a defined structure such as JSON, a grammar, or a schema. Related: Prefill, Decode, KV cache, PagedAttention. ### Tail latency Anchor: https://aruntime.com/glossary/#tail-latency Category: Observability and evaluation; layer: Observability and evaluation. Latency near the high end of a distribution, such as the 95th or 99th percentile, which exposes slow-path and queueing behavior hidden by averages. Related: Runtime trace, Span, Trace correlation, Replay. ### Tenant isolation Anchor: https://aruntime.com/glossary/#tenant-isolation Category: Security, policy, and evidence; layer: Security, policy, and evidence. Controls that prevent one tenant’s data, cache, memory, configuration, tools, or evidence from becoming accessible to another tenant. Related: Prompt injection, Confused deputy, Policy decision point, Policy enforcement point. ### Tensor parallelism Anchor: https://aruntime.com/glossary/#tensor-parallelism Category: Serving and distributed execution; layer: Serving and distributed execution. Distribution of tensor operations or parameter dimensions for one model execution across multiple devices. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### Thermal throttling Anchor: https://aruntime.com/glossary/#thermal-throttling Category: Edge, browser, and local; layer: Edge, browser, and local. Automatic reduction of device performance to remain within temperature or power limits. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### Throughput Anchor: https://aruntime.com/glossary/#throughput Category: Observability and evaluation; layer: Observability and evaluation. The amount of completed work per unit time, such as requests, tokens, or model executions, with workload and success criteria stated. Related: Runtime trace, Span, Trace correlation, Replay. ### TinyML Anchor: https://aruntime.com/glossary/#tinyml Category: Edge, browser, and local; layer: Edge, browser, and local. Deployment of machine-learning inference on highly resource-constrained microcontrollers or embedded systems. Related: Local runtime, Edge runtime, On-device runtime, Browser runtime. ### Tool broker Anchor: https://aruntime.com/glossary/#tool-broker Category: Agentic and application runtime; layer: Agentic and application runtime. A controlled component that resolves tool contracts, authorizes calls, validates arguments, invokes implementations, and records outcomes and side effects. Related: Agentic runtime, Agent framework, Tool contract, Side effect. ### Tool contract Anchor: https://aruntime.com/glossary/#tool-contract Category: Agentic and application runtime; layer: Agentic and application runtime. A versioned specification for a tool’s identifier, schemas, authentication, permissions, side effects, egress, timeout, retry, idempotency, approvals, compensation, errors, and evidence. Related: Agentic runtime, Agent framework, Tool broker, Side effect. ### TOSA Anchor: https://aruntime.com/glossary/#tosa Category: Compiler and graph; layer: Compiler and graph. A tensor-operator specification designed as an intermediate layer for machine-learning compilation and deployment. Related: Intermediate representation, ONNX, MLIR, StableHLO. ### TPOT Anchor: https://aruntime.com/glossary/#tpot Category: Inference and memory; layer: Inference and memory. Time per output token: an aggregate interval between generated output tokens, excluding or including specified phases according to the benchmark method. Related: Prefill, Decode, KV cache, PagedAttention. ### Trace correlation Anchor: https://aruntime.com/glossary/#trace-correlation Category: Observability and evaluation; layer: Observability and evaluation. Use of identifiers and links to connect infrastructure, model, tool, policy, product-outcome, and evaluation records without requiring all payloads in one system. Related: Runtime trace, Span, Replay, Tail latency. ### TTFT Anchor: https://aruntime.com/glossary/#ttft Category: Inference and memory; layer: Inference and memory. Time to first token: elapsed time from request admission or arrival to the first generated token, with the measurement boundary stated explicitly. Related: Prefill, Decode, KV cache, PagedAttention. ### Warmup Anchor: https://aruntime.com/glossary/#warmup Category: Serving and distributed execution; layer: Serving and distributed execution. Controlled execution used to initialize model state, compile paths, allocate memory, populate caches, or stabilize measurements before serving or benchmarking. Related: Dynamic batching, Admission control, Backpressure, Goodput. ### WebAssembly Anchor: https://aruntime.com/glossary/#webassembly Category: Edge, browser, and local; layer: Edge, browser, and local. A portable binary instruction format and sandboxed execution model used by browsers and other hosts for near-native compiled code. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### WebGPU Anchor: https://aruntime.com/glossary/#webgpu Category: Edge, browser, and local; layer: Edge, browser, and local. A web API exposing modern GPU rendering and general-purpose compute capabilities through a browser-managed security model. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### WebNN Anchor: https://aruntime.com/glossary/#webnn Category: Edge, browser, and local; layer: Edge, browser, and local. A web API for constructing and executing neural-network graphs using available operating-system and device acceleration capabilities. Related: Local runtime, Edge runtime, On-device runtime, TinyML. ### Working memory Anchor: https://aruntime.com/glossary/#working-memory Category: Agentic and application runtime; layer: Agentic and application runtime. Short-lived state retained for an active request, task, or session and bounded by explicit access and retention rules. Related: Agentic runtime, Agent framework, Tool broker, Tool contract. ## Diagram set - Seven-layer AI runtime stack: The runtime stack separates hardware, kernels, compilation, inference, serving, agentic control, and product workflow responsibilities. Source ID: seven-layer-stack. - Model execution path: A model artifact becomes executable through import, intermediate representation, optimization, lowering, loading, and hardware execution. Source ID: model-execution-path. - Request execution path: A production request carries identity, authority, risk, context, model-route, tool, policy, evidence, and memory decisions. Source ID: request-execution-path. - Control plane and execution plane: The control plane defines desired state; the execution plane handles live requests under those versioned decisions. Source ID: control-execution-planes. - Agentic request lifecycle: A governed agentic runtime treats planning and tool use as one part of a larger admission-to-evidence lifecycle. Source ID: agentic-request-lifecycle. - Tool authorization sequence: Tool selection does not grant authority; the runtime resolves policy and approval before invoking a side effect. Source ID: tool-authorization-sequence. - Memory-boundary model: Working, session, durable, and external memory have different authority, retention, and data-minimization rules. Source ID: memory-boundary-model. - Evidence and trace lifecycle: Correlation joins infrastructure, model, tool, policy, business-outcome, and evaluation records without requiring raw prompt retention. Source ID: evidence-trace-lifecycle. - Failure and recovery state machine: Retries are allowed only at classified safe points; partial side effects require reconciliation or compensation. Source ID: failure-recovery-state-machine. - Deployment topology comparison: Hosted, cloud or data-center, local, browser, edge, serverless, and hybrid deployments move different runtime responsibilities across trust boundaries. Source ID: deployment-topologies. - Runtime category boundary map: Categories describe responsibilities, not marketing labels; adjacent products can overlap without becoming interchangeable. Source ID: runtime-category-boundaries. - A product spanning multiple runtime layers: An illustrative platform can implement several layers while each responsibility remains separately testable and replaceable. Source ID: multi-layer-product. ## Developer artifacts - Runtime request schema: https://aruntime.com/runtime-request-contract/ and schemas/runtime-request.v1.schema.json - Tool contract schema: https://aruntime.com/tool-contracts/ and schemas/tool-contract.v1.schema.json - Policy and approval schema: https://aruntime.com/policy-approval-contracts/ and schemas/policy-approval.v1.schema.json - Evidence schema: https://aruntime.com/evidence-schema/ and schemas/evidence-record.v1.schema.json - Trace schema: https://aruntime.com/trace-schema/ and schemas/trace-envelope.v1.schema.json - Runnable PHP example: examples/php/runtime_pipeline.php ## Editorial controls ARuntime.com distinguishes sourced facts, editorial synthesis, and explicit proposals. Product capabilities are scoped to cited official documentation or maintainer repositories, include UTC verification dates, and are not converted into a universal score. The taxonomy is an editorial framework rather than a formal standard. Correction and material-change records are public site features.