Vai al contenuto
AI.info

tools

Braintrust

Braintrust helps teams trace AI applications, evaluate outputs, find production issues, and improve agents.

Braintrust

In inglese

Braintrust is an AI observability and evaluation platform for engineering, product, and AI teams. It records agent traces, prompts, responses, tool calls, latency, cost, and quality data.

Teams can run evaluations against datasets, compare prompts and models, score outputs with code, models, or humans, investigate recurring behaviors, and turn production traces into regression datasets. It does not provide a general-purpose AI model; model usage and some data processing or scoring incur additional charges.

Features

  • Inspect prompts, responses, and tool calls in real time
  • Run evaluations against versioned datasets
  • Compare prompts and models side by side
  • Score outputs with LLMs, code, or human reviewers
  • Search production traces and identify recurring behaviors
  • Turn production traces into evaluation datasets
  • Query logs and run evaluations through its MCP server
  • Use SDKs for Python, TypeScript, Go, Ruby, C#, and more

Use cases

  • Trace production AI agents and tool calls
  • Evaluate customer-support responses before release
  • Compare prompts and models against test datasets
  • Find recurring failures across production traces
  • Build regression tests from real failures
  • Monitor latency, cost, and response quality

Pros

    Cons

      Latest updates

      • September 2026 (v0.21.0+)

        Braintrust-hosted deployments map OpenInference traces to structured data; the bt CLI supports local JavaScript span plugins.

      • August 2026

        Monitoring views became dashboards; permissions now control access to organization AI providers; coding-agent tracing runs through the bt CLI.

      • July 2026

        Experiment comparisons support pairwise scoring; datasets can store references to full traces and show them in a Trace viewer.

      • June 2026

        SQL queries can run asynchronously through the API or bt sql --async, with results available for paginated retrieval.

      • May 2026

        Topics pricing adds monthly credits; custom facets are available on all plans, and trace timelines show cached-token usage.

      Capabilities

      • Choice of models — “compare prompts and models side-by-side” source
      • Self-hosted — “Deploy Brainstore data plane on your own infrastructure” source
      • API — “You get a unified API across OpenAI, Anthropic, Google, AWS, and other providers, with automatic caching and observability on every request.” source
      • Official SDKs — “Native SDKs SDKs for Python, TypeScript, Go, Ruby, C#, and more.” source
      • Traces and evaluates — “Full visibility into every agent interaction.” source

      Get it

      Security

      • SOC 2 Type II — “SOC 2 Type II - May 5 - Aug 5, 2025” source
      • SOC 2 — “SOC 2 Type II - May 5 - Aug 5, 2025” source
      • GDPR — “Data Processing Agreement (DPA) - Enterprise” source
      • HIPAA — “One-page white paper on Braintrust and HIPAA compliance.” source

      Pricing

      Starting price
      $0.50/mo
      Prices checked
      2026-09-26

      Starter

      • $0 / month
      • + $4/GB
      • + $2.50/1k
      • $10 model credits/month included
      • 1 GB processed data/month included
      • 10K scores/month included
      • 14-day retention
      • Unlimited users, projects, datasets, playgrounds, and experiments

      Pro

      • $249 / month
      • + $3/GB
      • + $1.50/1k
      • + $0.50/GB/mo
      • $100 model credits/month included
      • 5 GB processed data/month included
      • 50K scores/month included
      • 30-day retention
      • Custom charts, environments, priority support, and RBAC
      • Unlimited users, projects, datasets, playgrounds, and experiments

      Enterprise

      Price on request

      • Custom data retention and export
      • RBAC and premium support
      • On-prem or hosted deployment for high-volume or privacy-sensitive data
      Official website