Search ARuntime.com

Find runtime definitions and implementation guidance

Search page titles, summaries, headings, glossary terms, use cases, and runtime-directory entries.

Enter at least two characters.

Directory

Runtime Directory

Browse a vendor-neutral AI runtime directory covering compiler and graph runtimes, inference engines, model servers, serving platforms, edge, browser, and local runtime systems.

Audience: Architects and developers evaluating runtime products Reading time: 8 minutes Status: Research Last reviewed:

Key takeaways

  • Directory categories describe primary responsibility; many products span more than one layer.
  • Profiles retain official sources and UTC review dates rather than copying vendor marketing.
  • A model format, protocol, or compiler IR is listed only when its role is explicitly distinguished from a complete runtime.
  • Directory inclusion is not an endorsement, maturity claim, or performance ranking.
  • Use the comparison guide and selection guide before treating two entries as direct alternatives.

Runtime boundary

A useful architecture identifies what this layer receives, owns, emits, measures, and refuses to own. That boundary prevents overlapping products from being treated as interchangeable.

Receives

Official project documentation, repositories, standards, release information, category definitions, and reviewed profile metadata.

Owns

Classification, source provenance, review cadence, correction workflow, and neutral profile language.

Emits

Filterable profiles with category, stack layer, maintainer, capabilities, sources, review date, and related ARuntime.com guidance.

Does not own

Vendor certification, universal benchmarks, support guarantees, pricing, or a declaration that one system is best.

Failure modes

Stale project status, category confusion, copied marketing claims, broken official links, unscoped comparisons, and profile drift.

Evidence and metrics

Profiles reviewed, source quality, review age, broken links, correction time, category coverage, and unresolved claims.

How to use the directory

Filter by name, category, layer, maintainer, or capability, then open a profile and its official sources. Treat each profile as a starting point for a requirements-driven proof.

Implementation

Keep profile fields structured and versionable. Link names to canonical profile URLs and preserve source access dates.

Operational implications

Schedule review for time-sensitive features and project status. Route corrections through the public correction workflow.

Measure

Profile review age, source-link health, missing fields, and correction turnaround.

Category boundaries

Compiler runtimes, inference engines, model servers, serving platforms, local runtimes, edge/browser runtimes, and agentic infrastructure solve overlapping but distinct problems.

Implementation

Record primary and secondary layers, delegated backends, model formats, hardware targets, and external components.

Operational implications

Avoid a flat feature checklist across unlike categories. Use the taxonomy before comparison.

Measure

Category coverage, ambiguous profiles, external dependency count, and classification corrections.

Profile evidence

Profiles should prefer official documentation and repositories, retain a UTC review date, and qualify changing features or project status.

Implementation

Store source title, publisher, URL, type, access date, and page sections used.

Operational implications

Do not publish live pricing, performance, maintenance, or support claims without current verification and scope.

Measure

Primary-source share, source age, broken links, and unverified claims.

Runtime profiles

The filterable profile grid below is generated from the same structured PHP data used to seed individual runtime profile posts.

Implementation

Server-render the full directory and use lightweight vanilla JavaScript only for progressive filtering.

Operational implications

The directory remains usable without JavaScript and exposes shareable canonical profile URLs.

Measure

Rendered profile count, filter accessibility, keyboard behavior, and profile-link validity.

Filterable runtime profiles

Filter the reviewed profiles by runtime name, category, layer, maintainer, or capability. All profiles remain visible and usable without JavaScript.

Filter the runtime directory









Clear filters

16 profiles match the current scope.

Compiler and graph runtime

ONNX Runtime

Executes ONNX graphs through a common runtime API and hardware-specific execution providers.

Layers
Layer 2, Layer 3
Deployment
Cloud or data center, Local or desktop, Edge or mobile
Verified
2026-06-23T00:00:00Z

Model and LLM inference engine

vLLM

LLM inference and serving engine focused on efficient batching, memory management, and token generation.

Layers
Layer 3, Layer 4
Deployment
Cloud or data center, Local or desktop
Verified
2026-06-23T00:00:00Z

Model and LLM inference engine

SGLang

Inference and serving framework for structured language-model programs, prefix reuse, and efficient generation.

Layers
Layer 3, Layer 4
Deployment
Cloud or data center, Local or desktop
Verified
2026-06-23T00:00:00Z

Model serving runtime

NVIDIA Triton Inference Server

Model-serving runtime with model repositories, version handling, health endpoints, schedulers, batching, and multiple backends.

Layers
Layer 4
Deployment
Cloud or data center, Edge where supported
Verified
2026-06-23T00:00:00Z

Compiler and graph runtime

NVIDIA TensorRT

SDK that optimizes neural-network graphs and executes generated engines on supported NVIDIA hardware.

Layers
Layer 2, Layer 3
Deployment
Cloud or data center, Local or desktop, Edge on supported NVIDIA platforms
Verified
2026-06-23T00:00:00Z

Model and LLM inference engine

TensorRT-LLM

LLM inference stack that builds and executes TensorRT-based engines with model-specific and distributed optimizations.

Layers
Layer 2, Layer 3, Layer 4
Deployment
Cloud or data center, Local or desktop
Verified
2026-06-23T00:00:00Z

Compiler and graph runtime

Apache TVM

Machine-learning compiler stack for importing models, transforming IR, scheduling tensor programs, and generating target code.

Layers
Layer 1, Layer 2
Deployment
Cloud or data center, Local or desktop, Edge or mobile
Verified
2026-06-23T00:00:00Z

Intermediate representation

StableHLO

Portable high-level operation set used between machine-learning framework frontends and compiler backends.

Layers
Layer 2
Deployment
Build-time compiler pipeline
Verified
2026-06-23T00:00:00Z

Edge and mobile inference runtime

ExecuTorch

PyTorch on-device inference stack for mobile, embedded, and edge deployment.

Layers
Layer 2, Layer 3
Deployment
Edge or mobile, Local or desktop, Embedded
Verified
2026-06-23T00:00:00Z

Edge and mobile inference runtime

LiteRT

On-device runtime and tooling for executing optimized machine-learning models across supported mobile and edge hardware.

Layers
Layer 3
Deployment
Edge or mobile, Embedded, Local or desktop
Verified
2026-06-23T00:00:00Z

Browser inference API

WebNN

Web API for constructing and executing neural-network graphs through operating-system and hardware machine-learning capabilities.

Layers
Layer 2, Layer 3
Deployment
Browser
Verified
2026-06-23T00:00:00Z

Browser compute API

WebGPU

Web API exposing modern GPU computation used by browser inference libraries and other compute workloads.

Layers
Layer 1, Layer 2, Layer 3
Deployment
Browser, Local or desktop through web views
Verified
2026-06-23T00:00:00Z

Compiler and graph runtime

OpenVINO

Model optimization and inference runtime stack for supported Intel hardware targets.

Layers
Layer 2, Layer 3
Deployment
Cloud or data center, Local or desktop, Edge or mobile
Verified
2026-06-23T00:00:00Z

Serving platform

KServe

Kubernetes-native serving platform for deploying, scaling, and operating model inference services.

Layers
Layer 4
Deployment
Cloud or data center, Kubernetes
Verified
2026-06-23T00:00:00Z

Model serving runtime

Ray Serve

Distributed serving layer for composing and scaling Python model-backed applications on Ray.

Layers
Layer 4
Deployment
Cloud or data center, Local or desktop for development
Verified
2026-06-23T00:00:00Z

Model serving framework

BentoML

Framework for packaging model-backed Python services and exposing them through deployment-oriented APIs.

Layers
Layer 4
Deployment
Cloud or data center, Local or desktop for development
Verified
2026-06-23T00:00:00Z

Reference tables

Directory category guide
Category Primary responsibility Typical input Typical output What it is not
Compiler / graph runtime Import, optimize, partition, lower, and execute graphs Framework graph or portable model Optimized graph, executable plan, predictions A complete serving platform
LLM inference engine Efficient prefill, decode, KV cache, batching, and streaming LLM weights and token requests Generated token stream A complete agent runtime
Model server Expose model execution through APIs and lifecycle controls Network request and model repository Prediction/stream plus server telemetry A Kubernetes control plane
Serving platform Deploy, scale, route, and roll out model services Runtime definition and deployment spec Managed inference service The model kernel itself
Edge/mobile runtime Execute prepared models on constrained devices AOT model program and local input On-device prediction A universal cloud service
Browser runtime/API Execute model graphs or kernels in a web sandbox Web assets and client input Client-local prediction Guaranteed support on every browser
Agentic runtime infrastructure Coordinate context, tools, memory, policy, durability, and evaluation Governed task envelope Traceable task outcome Only a prompt or tool-calling library

Decision checklist

  1. Which runtime layer and responsibility does the reader need?
  2. Is the candidate a complete product, a backend, a format, a protocol, or a compiler component?
  3. Which official source verifies the feature or status being evaluated?
  4. What model, hardware, deployment, and operating requirements must the candidate satisfy?
  5. Which external components are required to form a production stack?
  6. What proof and review date are required before selection?

Common mistakes

  • Treating directory order as a ranking.
  • Comparing a model server directly with a graph compiler without a scope statement.
  • Copying unqualified vendor performance claims.
  • Calling a model format or protocol a complete runtime.
  • Leaving project status and feature claims undated.
  • Assuming a listed capability is enabled by default for every model and hardware target.

Sources and further reading


  1. ONNX Runtime high-level design
    (opens in a new tab)

    ONNX Runtime · Official documentation · reviewed 2026-06-21 UTC

  2. vLLM documentation
    (opens in a new tab)

    vLLM · Official documentation · reviewed 2026-06-21 UTC

  3. Triton Inference Server architecture
    (opens in a new tab)

    NVIDIA · Official documentation · reviewed 2026-06-21 UTC

  4. KServe ServingRuntime
    (opens in a new tab)

    KServe · Official documentation · reviewed 2026-06-21 UTC

  5. ExecuTorch overview
    (opens in a new tab)

    PyTorch · Official documentation · reviewed 2026-06-21 UTC

Last reviewed: 2026-06-23 UTC

Maintenance record

Found an error, outdated capability, or unclear category boundary? Submit a correction with a supporting source.