tools
Weights & Biases - ML Experiment Tracking & LLM Evaluation | AI.info
Weights & Biases is the standard MLOps platform for experiment tracking, model versioning, and LLM evaluation. Used by 500,000+ ML practitioners at OpenAI, NVIDIA, and beyond.

In inglese
Weights & Biases helps machine learning and AI teams track experiments, compare model runs, version datasets and models, and share results. Its platform includes W&B Models for model development and W&B Weave for tracing, evaluating, and monitoring AI applications and agents.
It is used by researchers, ML engineers, AI developers, and enterprise teams. The platform offers cloud-hosted and self-hosted deployment options, while inference, storage, data ingestion, and some sandbox usage can incur additional charges.
Features
- Track and visualize machine learning experiments
- Run automated hyperparameter sweeps
- Version datasets and models with W&B Artifacts
- Manage models, datasets, prompts, code, and metadata in a registry
- Trace and debug AI applications with W&B Weave
- Evaluate AI applications and compare model or prompt performance
- Monitor production AI applications and agent workflows
- Share findings through reports, tables, and collaborative dashboards
Use cases
- Track training runs and compare model performance
- Tune hyperparameters across distributed experiments
- Version datasets and model checkpoints for reproducibility
- Evaluate prompts, RAG pipelines, and agent applications
- Monitor production AI quality, latency, cost, and safety
- Coordinate model development across research and engineering teams
Pros
Cons
Capabilities
- Command line — “You can then use the command line to start the server.” source
- Self-hosted — “Self-hosting allows you to run Weights & Biases on your own infrastructure, giving you full control over your data and privacy.” source
- API — “Each inference API has different costs for input and output tokens.” source
- Official SDKs — “Core | Reports | Automations | SDK” source
- Runs models for you — “Inference Access and explore hosted AI models” source
- Builds agents and workflows — “W&B Weave: Build agentic AI applications” source
- Traces and evaluates — “Monitor and analyze the performance of your GenAI models during development and in production, capturing inputs, outputs, and metadata for each inference.” source
Get it
Security
- SOC 2 Type II — “We are certified under ISO/IEC 27001:2022, ISO/IEC 27017:2015, and ISO/IEC 27018:2019, and remain compliant with SOC 2 Type 2 and HIPAA standards.” source
- SOC 2 — “We are certified under ISO/IEC 27001:2022, ISO/IEC 27017:2015, and ISO/IEC 27018:2019, and remain compliant with SOC 2 Type 2 and HIPAA standards.” source
- ISO 27001 — “We are certified under ISO/IEC 27001:2022, ISO/IEC 27017:2015, and ISO/IEC 27018:2019, and remain compliant with SOC 2 Type 2 and HIPAA standards.” source
- HIPAA — “We are certified under ISO/IEC 27001:2022, ISO/IEC 27017:2015, and ISO/IEC 27018:2019, and remain compliant with SOC 2 Type 2 and HIPAA standards.” source
Pricing
- Starting price
- $60/mo
- Prices checked
- 2026-09-25
Free
Free
- AI application evaluations
- AI application tracing
- AI application scorers
- AI model experiment tracking
- AI assets registry & lineage tracking
- Community Support
Pro
- Starts at $60/month, billed monthly
- Unlimited teams for collaboration
- Team-based access controls
- Service Accounts
- Priority email & chat support
- CI/CD automations
- Slack and email alerts
Enterprise
Price on request
- Single tenant option with choice of region
- HIPAA compliant option
- Secure private connectivity
- Customer-managed encryption key
- Single Sign On
- Automated user provisioning
Personal
Free
- 1 user seat
- Experiment tracking
- Registry & lineage tracking
- Run a W&B server locally on any machine with Docker and Python installed
- For personal projects only. Corporate use is not allowed.
Advanced Enterprise
Price on request
- Flexible deployment options
- HIPAA compliant option
- Secure private connectivity
- Customer-managed encryption key
- Single Sign On
- Automated user provisioning
Free forever for academic research
Free
- All the product features included on Pro
- Unlimited tracked hours
- 200GB of cloud storage
- Up to 25GB/mo of Weave data ingestion
- Up to 100 seats
- Additional cloud storage can be purchased for $0.03 per GB, billed monthly