tools
DeepEval
Open-source framework for testing and evaluating LLM apps, agents, RAG pipelines, chatbots, and multimodal systems.

In inglese
DeepEval runs LLM evaluations locally through Python or TypeScript, with Pytest-style assertions, 50+ metrics, tracing, synthetic dataset generation, conversation simulation, and component-level testing. It supports agents, RAG pipelines, chatbots, MCP systems, and custom workflows.
Developers use the CLI in local development and CI/CD, including with coding agents such as Cursor, Claude Code, and Codex. Shared dashboards, regression tracking, observability, production monitoring, and other hosted features are provided through the separate Confident AI platform and may cost extra.
Features
- Run LLM evaluations locally with Pytest-style assertions
- Use 50+ metrics for agents, RAG, safety, conversations, and multimodal apps
- Trace agent runs and evaluate complete trajectories or individual components
- Generate synthetic datasets and edge-case goldens
- Run evaluation suites from a CLI in local development or CI/CD
- Support Python and TypeScript SDKs
- Integrate with LangChain, LangGraph, OpenAI Agents, and OpenTelemetry
- Apache 2.0 licensed open-source framework
Use cases
- Test chatbot answers for relevance, correctness, safety, and faithfulness
- Evaluate RAG retrieval and generated responses against reference data
- Grade agent trajectories, tool calls, and component-level behavior
- Run regression tests on every pull request in CI/CD
- Generate synthetic test datasets from documents or application code
- Inspect local traces while iterating with a coding agent
Pros
Cons
Latest updates
- typescript-v0.9.20 (v0.9.20)
New release.
- typescript-v0.9.19 (v0.9.19)
New release.
- python-v4.2.6 (v4.2.6)
New release.
- typescript-v0.9.17 (v0.9.17)
New release.
- Introducing Jev in DeepEval (4.2.2)
DeepEval 4.2.2 introduces support for Jev, TypeSafe AI's System One model.
Capabilities
- Runs commands — “DeepEval traces every step of your agent into something you can grade, and improve — visible in your terminal, testable in your runner, shippable in your next commit.” source
- Command line — “Cursor, Claude Code, and Codex shell out to one CLI, read scored traces with reasons, then patch the failing span and re-run to confirm.” source
- Self-hosted — “Iterate locally, on your own environment, on your own criteria.” source
- API — “Pytest-native evals that run in CI/CD or as Python scripts.” source
- Traces and evaluates — “DeepEval integrates natively with Confident AI, an AI observability and evaluation platform.” source
Get it
Pricing
- Prices checked
- 2026-09-26