Tools
Baseten vs Fireworks AI
Baseten or Fireworks AI? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In inglese
In short
- Both: Choice of models, API, Runs models for you
- Only Baseten states: Self-hosted
- Only Fireworks AI states: Command line, Official SDKs
Baseten
Baseten provides managed infrastructure to deploy, serve, train, and distribute open-source, custom, and fine-tuned AI models.
Plans
Basic
Free
- Dedicated deployments
- Model APIs
- Training
- Fast cold starts
- SOC 2 Type II and HIPAA compliant
- Email and in-app chat support
Dedicated Deployments — T4
- $0.01052 per minute
- 16 GiB VM
Dedicated Deployments — L4
- $0.01414 per minute
- 24 GiB VRAM
Dedicated Deployments — A10G
- $0.02012 per minute
- 24 GiB VM
Dedicated Deployments — A100
- $0.06667 per minute
- 80 GiB VRAM
Dedicated Deployments — H100 MIG
- $0.0625 per minute
- 40 GiB VRAM
Dedicated Deployments — H100
- $0.10833 per minute
- 80 GiB VRAM
Dedicated Deployments — B200
- $0.16633 per minute
- 180 GiB VRAM
Prices checked 2026-09-24 on the maker’s page.
Capabilities
- Choice of models — “You can deploy open source and custom models on Baseten.” source
- Self-hosted — “Yes, you can self-host Baseten in order to manage security and use your own cloud commitments.” source
- API — “Model APIs” source
- Runs models for you — “Instant access to pre-optimized models running on the Baseten Inference Stack.” source
Security
- SOC 2 Type II — “Baseten Labs, Inc. SOC 2 Type 2 Report 5.31.26.pdf” source
- SOC 2 — “Baseten Labs, Inc. SOC 2 Type 2 Report 5.31.26.pdf” source
- ISO 27001 — “Baseten Labs, Inc. ISO 27001 Certificate.pdf” source
- GDPR — “SOC 2 ISO 27001:2022 HIPAA CCPA GDPR PCI DSS - SAQ D” source
- HIPAA — “Baseten Labs, Inc. HIPAA Report 5.31.26.pdf” source
Latest updates
- Web search with Baseten Hosted Tools
Baseten Hosted Tools now brings server-side web search to Model APIs through Baseten Grounded Inference.
- Baseten CLI 1.0.0 (1.0.0)
The Baseten CLI is now generally available. Version 1.0.0 marks the command surface as stable.
- Model API Deprecation (GLM 4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling Small, DeepSeek v4 Pro)
GLM 4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling Small, and DeepSeek v4 Pro will be deprecated at 5pm PT September 25th.
Fireworks AI
Fireworks AI hosts, fine-tunes, and serves open models through serverless and dedicated APIs.
Plans
Serverless Inference — Embeddings (up to 150M parameters)
- $0.008 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Embeddings (150M–350M parameters)
- $0.016 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Qwen3 8B
- $0.1 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Managed Training — Models up to 16B parameters
- $0.50 / 1M training tokens
- $1.00 / 1M training tokens
- $2.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 16.1B–80B
- $3.00 / 1M training tokens
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 80B–300B
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- $24.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models over 300B
- $10.00 / 1M training tokens
- $20.00 / 1M training tokens
- $40.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Serverless Training API — GLM 5.3
- $4.86 / 1M Prefill
- $0.972 / 1M Cached Prefill
- $12.15 / 1M Sample
- $14.58 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Qwen 3.8 27B
- $1.86 / 1M Prefill
- $0.372 / 1M Cached Prefill
- $5.595 / 1M Sample
- $4.103 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Kimi K3
- $10.87 / 1M Prefill
- $2.17 / 1M Cached Prefill
- $27.11 / 1M Sample
- $32.55 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Command line — “Developers Model Library Docs CLI API Changelog” source
- Choice of models — “Route to the best open or closed model for every task, and cut your AI coding spend 50 to 75%.” source
- API — “Serverless. Pay per token with Priority and Fast options to meet your requirements. OpenAI and Anthropic compatible.” source
- Official SDKs — “The Fireworks Training SDK lets us focus on our research instead of wrestling with infrastructure.” source
- Runs models for you — “Serve the latest open models, or your own trained versions.” source
Security
- SOC 2 Type II — “SOC 2 Type 2” source
- SOC 2 — “SOC 2 Type 2” source
- ISO 27001 — “ISO 27001 Certificate” source
- ISO 42001 — “ISO 42001 Certificate” source
- GDPR — “Compliance SOC 2 Type 2 HIPAA GDPR” source
- HIPAA — “SOC 2 Type 2 HIPAA” source
Latest updates
- Serverless pricing update: DeepSeek V4.1 Flash
Serverless pricing for DeepSeek V4.1 Flash changes; dedicated deployment and Reserved Throughput pricing is unaffected.
- New deployment creation flags: deploymentShape: "default" and acceptShapelessRisk
Create Deployment adds deploymentShape: "default" to pick a validated deployment shape and acceptShapelessRisk to create without a shape.
- Upcoming Serverless deprecation: older DeepSeek, GLM, Muse, and Kimi models
Several older Serverless models will be decommissioned on September 25, 2026; migrate to a recommended replacement before then.