tools
Baseten
Baseten provides managed infrastructure to deploy, serve, train, and distribute open-source, custom, and fine-tuned AI models.

In inglese
Baseten lets teams deploy custom and open-source models, use pre-optimized Model APIs, run training jobs, and scale inference on managed GPU infrastructure. It supports OpenAI-compatible APIs, dedicated deployments, autoscaling, observability, and self-hosted deployments for enterprise customers.
Teams pay for compute or Model API usage. The Basic plan has no monthly platform fee but uses pay-as-you-go pricing; Pro and Enterprise pricing requires a quote. Baseten does not provide a general-purpose end-user chatbot.
Features
- Deploy open-source, custom, and fine-tuned models on dedicated infrastructure
- Use pre-optimized Model APIs with OpenAI-compatible endpoints
- Scale deployments with autoscaling and configurable GPU resources
- Run multi-node training jobs with checkpoint syncing
- Train with the Loops SDK and deploy checkpoints to inference
- Monitor deployments with logs, metrics, and observability APIs
- Use server-side hosted tools such as web search
- Package and serve models with the open-source Truss framework
Use cases
- Deploy custom models as production inference endpoints
- Integrate open-source Model APIs into an application
- Train and deploy models using managed GPU infrastructure
- Scale high-volume inference workloads across GPU instances
- Monitor model latency, logs, metrics, and usage
- Distribute models to customers through Baseten for Model Labs
Pros
Cons
Latest updates
- Web search with Baseten Hosted Tools
Baseten Hosted Tools now brings server-side web search to Model APIs through Baseten Grounded Inference.
- Baseten CLI 1.0.0 (1.0.0)
The Baseten CLI is now generally available. Version 1.0.0 marks the command surface as stable.
- Model API Deprecation (GLM 4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling Small, DeepSeek v4 Pro)
GLM 4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling Small, and DeepSeek v4 Pro will be deprecated at 5pm PT September 25th.
- OIDC and AWS AssumeRole for training jobs
Training jobs can now use OIDC or AWS AssumeRole during setup to pull private container images and download model weights or training data without storing long-lived…
- Model API costs
GET /v1/billing/model_apis returns exact subtotals for each calendar day, rated from your usage.
Capabilities
- Choice of models — “You can deploy open source and custom models on Baseten.” source
- Self-hosted — “Yes, you can self-host Baseten in order to manage security and use your own cloud commitments.” source
- API — “Model APIs” source
- Runs models for you — “Instant access to pre-optimized models running on the Baseten Inference Stack.” source
Get it
Security
- SOC 2 Type II — “Baseten Labs, Inc. SOC 2 Type 2 Report 5.31.26.pdf” source
- SOC 2 — “Baseten Labs, Inc. SOC 2 Type 2 Report 5.31.26.pdf” source
- ISO 27001 — “Baseten Labs, Inc. ISO 27001 Certificate.pdf” source
- GDPR — “SOC 2 ISO 27001:2022 HIPAA CCPA GDPR PCI DSS - SAQ D” source
- HIPAA — “Baseten Labs, Inc. HIPAA Report 5.31.26.pdf” source
Pricing
- Starting price
- Free
- Prices checked
- 2026-09-24
Basic
Free
- Dedicated deployments
- Model APIs
- Training
- Fast cold starts
- SOC 2 Type II and HIPAA compliant
- Email and in-app chat support
Dedicated Deployments — T4
- $0.01052 per minute
- 16 GiB VM
Dedicated Deployments — L4
- $0.01414 per minute
- 24 GiB VRAM
Dedicated Deployments — A10G
- $0.02012 per minute
- 24 GiB VM
Dedicated Deployments — A100
- $0.06667 per minute
- 80 GiB VRAM
Dedicated Deployments — H100 MIG
- $0.0625 per minute
- 40 GiB VRAM
Dedicated Deployments — H100
- $0.10833 per minute
- 80 GiB VRAM
Dedicated Deployments — B200
- $0.16633 per minute
- 180 GiB VRAM