The field guide to AI eval tools, written for engineers.
We test, compare, and rate the platforms teams use to measure LLMs and agents in production — observability, offline & online evals, prompt management, gateways, and red-teaming.
- 25
- companies reviewed
- Aug 21, 2026
- last updated
Featured companies
Braintrust
9.1Eval-driven dev platform combining traces, datasets, scorers, and a playground in one product.
Fiddler
7.2Enterprise ML governance platform extended to LLMs and generative AI, with audit-ready traces and in-environment evaluations.
Galileo
7.5Agent reliability platform with cheap, fast evaluators that can run on every request in production.
HUD
7.8Open-source platform for building RL environments and evals for computer-use agents — used by frontier labs, ships its own benchmarks.
Langfuse
8.4Open-source LLM observability with evals, prompt management, and best-in-class tracing.
LiteLLM
8.0Open-source Python SDK and proxy that translates requests across 100+ LLM providers into the OpenAI format.
Recent editorial
The best AI eval tools for CI/CD (2026)
Five tools for running LLM and agent evals in pull requests, reporting useful diffs, and blocking releases when quality regresses.
The best AI evaluation tools (2026)
Five AI evaluation platforms ranked by the work that matters in production: building datasets, scoring outputs, comparing changes, catching regressions, and learning from live traffic.
The best LLM tracing tools (2026)
Five tracing platforms ranked for debugging multi-step LLM and agent applications: span depth, search, OpenTelemetry support, cost attribution, and the path from failure to regression test.