tools
Braintrust
Braintrust helps teams trace AI applications, evaluate outputs, find production issues, and improve agents.

In inglese
Braintrust is an AI observability and evaluation platform for engineering, product, and AI teams. It records agent traces, prompts, responses, tool calls, latency, cost, and quality data.
Teams can run evaluations against datasets, compare prompts and models, score outputs with code, models, or humans, investigate recurring behaviors, and turn production traces into regression datasets. It does not provide a general-purpose AI model; model usage and some data processing or scoring incur additional charges.
Features
- Inspect prompts, responses, and tool calls in real time
- Run evaluations against versioned datasets
- Compare prompts and models side by side
- Score outputs with LLMs, code, or human reviewers
- Search production traces and identify recurring behaviors
- Turn production traces into evaluation datasets
- Query logs and run evaluations through its MCP server
- Use SDKs for Python, TypeScript, Go, Ruby, C#, and more
Use cases
- Trace production AI agents and tool calls
- Evaluate customer-support responses before release
- Compare prompts and models against test datasets
- Find recurring failures across production traces
- Build regression tests from real failures
- Monitor latency, cost, and response quality
Pros
Cons
Latest updates
- September 2026 (v0.21.0+)
Braintrust-hosted deployments map OpenInference traces to structured data; the bt CLI supports local JavaScript span plugins.
- August 2026
Monitoring views became dashboards; permissions now control access to organization AI providers; coding-agent tracing runs through the bt CLI.
- July 2026
Experiment comparisons support pairwise scoring; datasets can store references to full traces and show them in a Trace viewer.
- June 2026
SQL queries can run asynchronously through the API or bt sql --async, with results available for paginated retrieval.
- May 2026
Topics pricing adds monthly credits; custom facets are available on all plans, and trace timelines show cached-token usage.
Capabilities
- Choice of models — “compare prompts and models side-by-side” source
- Self-hosted — “Deploy Brainstore data plane on your own infrastructure” source
- API — “You get a unified API across OpenAI, Anthropic, Google, AWS, and other providers, with automatic caching and observability on every request.” source
- Official SDKs — “Native SDKs SDKs for Python, TypeScript, Go, Ruby, C#, and more.” source
- Traces and evaluates — “Full visibility into every agent interaction.” source
Get it
Security
Pricing
- Starting price
- $0.50/mo
- Prices checked
- 2026-09-26
Starter
- $0 / month
- + $4/GB
- + $2.50/1k
- $10 model credits/month included
- 1 GB processed data/month included
- 10K scores/month included
- 14-day retention
- Unlimited users, projects, datasets, playgrounds, and experiments
Pro
- $249 / month
- + $3/GB
- + $1.50/1k
- + $0.50/GB/mo
- $100 model credits/month included
- 5 GB processed data/month included
- 50K scores/month included
- 30-day retention
- Custom charts, environments, priority support, and RBAC
- Unlimited users, projects, datasets, playgrounds, and experiments
Enterprise
Price on request
- Custom data retention and export
- RBAC and premium support
- On-prem or hosted deployment for high-volume or privacy-sensitive data