Tools
Baseten vs Together AI
Baseten or Together AI? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In short
- Both: Choice of models, API, Runs models for you
- Only Baseten states: Self-hosted
- Only Together AI states: Runs commands, Builds agents and workflows, Traces and evaluates
Baseten
Baseten provides managed infrastructure to deploy, serve, train, and distribute open-source, custom, and fine-tuned AI models.
Plans
Basic
Free
- Dedicated deployments
- Model APIs
- Training
- Fast cold starts
- SOC 2 Type II and HIPAA compliant
- Email and in-app chat support
Dedicated Deployments — T4
- $0.01052 per minute
- 16 GiB VM
Dedicated Deployments — L4
- $0.01414 per minute
- 24 GiB VRAM
Dedicated Deployments — A10G
- $0.02012 per minute
- 24 GiB VM
Dedicated Deployments — A100
- $0.06667 per minute
- 80 GiB VRAM
Dedicated Deployments — H100 MIG
- $0.0625 per minute
- 40 GiB VRAM
Dedicated Deployments — H100
- $0.10833 per minute
- 80 GiB VRAM
Dedicated Deployments — B200
- $0.16633 per minute
- 180 GiB VRAM
Prices checked 2026-09-24 on the maker’s page.
Capabilities
- Choice of models — “You can deploy open source and custom models on Baseten.” source
- Self-hosted — “Yes, you can self-host Baseten in order to manage security and use your own cloud commitments.” source
- API — “Model APIs” source
- Runs models for you — “Instant access to pre-optimized models running on the Baseten Inference Stack.” source
Security
- SOC 2 Type II — “Baseten Labs, Inc. SOC 2 Type 2 Report 5.31.26.pdf” source
- SOC 2 — “Baseten Labs, Inc. SOC 2 Type 2 Report 5.31.26.pdf” source
- ISO 27001 — “Baseten Labs, Inc. ISO 27001 Certificate.pdf” source
- GDPR — “SOC 2 ISO 27001:2022 HIPAA CCPA GDPR PCI DSS - SAQ D” source
- HIPAA — “Baseten Labs, Inc. HIPAA Report 5.31.26.pdf” source
Latest updates
- Web search with Baseten Hosted Tools
Baseten Hosted Tools now brings server-side web search to Model APIs through Baseten Grounded Inference.
- Baseten CLI 1.0.0 (1.0.0)
The Baseten CLI is now generally available. Version 1.0.0 marks the command surface as stable.
- Model API Deprecation (GLM 4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling Small, DeepSeek v4 Pro)
GLM 4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling Small, and DeepSeek v4 Pro will be deprecated at 5pm PT September 25th.
Together AI
Together AI provides APIs and infrastructure for running, fine-tuning, and training open AI models.
Plans
Serverless Inference
- MiniMax M3 — $0.30 per 1M tokens (Input)
- MiniMax M3 — $1.20 per 1M tokens (output)
- Kimi K3 — $3.00 per 1M tokens (Input)
- Kimi K3 — $15.00 per 1M tokens (output)
- GLM-5.3-Flash — $0.15 per 1M tokens (Input)
- GLM-5.3-Flash — $0.50 per 1M tokens (output)
- GPT Image 2 — $0.053 per image
- Wan 2.6 Image — $0.03 per image
- ByteDance Seedance 2.5 — $0.115 per video
- ByteDance Seedance 2.0 — $0.16 per video
- NVIDIA Nemotron 3 ASR Streaming 0.6B — $0.0015 per audio minute
- Whisper Large v3 — $0.0015 per audio minute
- High-performance inference as APIs
- Prices vary by model and task; the page also lists batch API prices.
Dedicated Inference — NVIDIA HGX H100
- $5.49 per gpu per hour
- $3.99 per gpu per hour
- Single-tenant GPU instances
- Guaranteed performance (no sharing)
- Support for custom models
- Autoscaling & traffic spike handling
Dedicated Inference — NVIDIA HGX B200
- $8.99 per gpu per hour
- Single-tenant GPU instances
- Guaranteed performance (no sharing)
- Support for custom models
- Autoscaling & traffic spike handling
Dedicated Inference — other hardware
Price on request
- NVIDIA HGX H200, NVIDIA HGX B300, NVIDIA GB200 NVL72, and NVIDIA GB300 NVL72
- Contact sales
GPU Clusters — On-demand
- NVIDIA HGX B200 $8.19 per GPU per hour
- NVIDIA HGX B300 $9.99 per GPU per hour
- NVIDIA HGX H100 $3.99 per GPU per hour
- NVIDIA HGX H200 $5.99 per GPU per hour
- Pay-as-you-go GPU capacity on an hourly basis
GPU Clusters — Preemptible and reserved
- NVIDIA HGX H100 Preemptible Compute $1.99 per GPU per hour
- NVIDIA HGX H100 ON-Demand $3.99 per GPU per hour
- NVIDIA HGX H100 7-30 days $3.69 per GPU per hour
- NVIDIA HGX H100 31-90 days $3.45 per GPU per hour
- NVIDIA HGX H100 91-180 days $3.19 per GPU per hour
- On-demand hourly rates and reserved capacity
- Reservation terms shown as 7-30, 31-90, and 91-180 days, and 181+ days
Code Sandbox
- Per vCPU $0.0446 per hour
- Per GiB RAM $0.0149 per hour
- Customize a deployment of VM sandboxes for large development environments
Code Interpreter
- Session (60 minutes) $0.03 per session
- Execute LLM-generated code securely using the API
Managed Storage
- Shared Filesystem $0.16 GiB/month
- High-bandwidth, parallel filesystem colocated with your compute
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Runs commands — “await client.commands.run("npm install && npm run build")” source
- Choice of models — “Scale to 30 billion tokens per model with any serverless model or private deployment.” source
- API — “High-performance inference as APIs” source
- Runs models for you — “The fastest way to run open-source models on demand.” source
- Builds agents and workflows — “Build voice agents for production” source
- Traces and evaluates — “Measure model quality” source
Latest updates
- How to train your own Jev for $17
Launched the together/Tev1-4B-experimental classifier on Together’s serverless platform.
- Canary rollouts: upgrade models in production without downtime
Dedicated inference supports staged traffic ramps, metric gates, and automatic rollback for model upgrades.
- Together AI expands fine-tuning service with more models, live metrics, and finer controls
Together Fine-Tuning added more models, live experiment tracking, Expert LoRA, early stopping, dataset previews, and pre-flight validation.