Search ARuntime.com

Find runtime definitions and implementation guidance

Search page titles, summaries, headings, glossary terms, use cases, and runtime-directory entries.

Enter at least two characters.

Build and Operate

Evaluation

Versioned offline and online evaluation for model behavior, runtime controls, tool use, policy, and business outcomes.

Audience: ML engineers; platform teams; product and governance leaders Reading time: 1 minute Status: Production guidance Last reviewed:

Evaluation determines whether a runtime configuration, model route, prompt or instruction version, retrieval policy, tool policy, and workflow produce acceptable outcomes under representative conditions.

Evaluation layers

  • Infrastructure: latency, errors, resource use, queueing, cache, and availability.
  • Model: task quality, structured-output validity, refusal behavior, and regression.
  • Tool and policy: permission decisions, argument validity, side effects, approval, and recovery.
  • Product: task completion, escalation, customer impact, and cost per successful outcome.

Release evidence

Version fixtures, expected outputs, graders, thresholds, model routes, policies, and datasets. Record known limitations and do not collapse safety, quality, cost, and latency into one universal score.

Maintenance record

Found an error, outdated capability, or unclear category boundary? Submit a correction with a supporting source.