Tools
Fireworks AI vs Together AI
Fireworks AI or Together AI? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In short
- Both: Choice of models, API, Runs models for you
- Only Fireworks AI states: Command line, Official SDKs
- Only Together AI states: Runs commands, Builds agents and workflows, Traces and evaluates
Fireworks AI
Fireworks AI hosts, fine-tunes, and serves open models through serverless and dedicated APIs.
Plans
Serverless Inference — Embeddings (up to 150M parameters)
- $0.008 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Embeddings (150M–350M parameters)
- $0.016 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Qwen3 8B
- $0.1 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Managed Training — Models up to 16B parameters
- $0.50 / 1M training tokens
- $1.00 / 1M training tokens
- $2.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 16.1B–80B
- $3.00 / 1M training tokens
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 80B–300B
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- $24.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models over 300B
- $10.00 / 1M training tokens
- $20.00 / 1M training tokens
- $40.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Serverless Training API — GLM 5.3
- $4.86 / 1M Prefill
- $0.972 / 1M Cached Prefill
- $12.15 / 1M Sample
- $14.58 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Qwen 3.8 27B
- $1.86 / 1M Prefill
- $0.372 / 1M Cached Prefill
- $5.595 / 1M Sample
- $4.103 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Kimi K3
- $10.87 / 1M Prefill
- $2.17 / 1M Cached Prefill
- $27.11 / 1M Sample
- $32.55 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Command line — “Developers Model Library Docs CLI API Changelog” source
- Choice of models — “Route to the best open or closed model for every task, and cut your AI coding spend 50 to 75%.” source
- API — “Serverless. Pay per token with Priority and Fast options to meet your requirements. OpenAI and Anthropic compatible.” source
- Official SDKs — “The Fireworks Training SDK lets us focus on our research instead of wrestling with infrastructure.” source
- Runs models for you — “Serve the latest open models, or your own trained versions.” source
Security
- SOC 2 Type II — “SOC 2 Type 2” source
- SOC 2 — “SOC 2 Type 2” source
- ISO 27001 — “ISO 27001 Certificate” source
- ISO 42001 — “ISO 42001 Certificate” source
- GDPR — “Compliance SOC 2 Type 2 HIPAA GDPR” source
- HIPAA — “SOC 2 Type 2 HIPAA” source
Latest updates
- Serverless pricing update: DeepSeek V4.1 Flash
Serverless pricing for DeepSeek V4.1 Flash changes; dedicated deployment and Reserved Throughput pricing is unaffected.
- New deployment creation flags: deploymentShape: "default" and acceptShapelessRisk
Create Deployment adds deploymentShape: "default" to pick a validated deployment shape and acceptShapelessRisk to create without a shape.
- Upcoming Serverless deprecation: older DeepSeek, GLM, Muse, and Kimi models
Several older Serverless models will be decommissioned on September 25, 2026; migrate to a recommended replacement before then.
Together AI
Together AI provides APIs and infrastructure for running, fine-tuning, and training open AI models.
Plans
Serverless Inference
- MiniMax M3 — $0.30 per 1M tokens (Input)
- MiniMax M3 — $1.20 per 1M tokens (output)
- Kimi K3 — $3.00 per 1M tokens (Input)
- Kimi K3 — $15.00 per 1M tokens (output)
- GLM-5.3-Flash — $0.15 per 1M tokens (Input)
- GLM-5.3-Flash — $0.50 per 1M tokens (output)
- GPT Image 2 — $0.053 per image
- Wan 2.6 Image — $0.03 per image
- ByteDance Seedance 2.5 — $0.115 per video
- ByteDance Seedance 2.0 — $0.16 per video
- NVIDIA Nemotron 3 ASR Streaming 0.6B — $0.0015 per audio minute
- Whisper Large v3 — $0.0015 per audio minute
- High-performance inference as APIs
- Prices vary by model and task; the page also lists batch API prices.
Dedicated Inference — NVIDIA HGX H100
- $5.49 per gpu per hour
- $3.99 per gpu per hour
- Single-tenant GPU instances
- Guaranteed performance (no sharing)
- Support for custom models
- Autoscaling & traffic spike handling
Dedicated Inference — NVIDIA HGX B200
- $8.99 per gpu per hour
- Single-tenant GPU instances
- Guaranteed performance (no sharing)
- Support for custom models
- Autoscaling & traffic spike handling
Dedicated Inference — other hardware
Price on request
- NVIDIA HGX H200, NVIDIA HGX B300, NVIDIA GB200 NVL72, and NVIDIA GB300 NVL72
- Contact sales
GPU Clusters — On-demand
- NVIDIA HGX B200 $8.19 per GPU per hour
- NVIDIA HGX B300 $9.99 per GPU per hour
- NVIDIA HGX H100 $3.99 per GPU per hour
- NVIDIA HGX H200 $5.99 per GPU per hour
- Pay-as-you-go GPU capacity on an hourly basis
GPU Clusters — Preemptible and reserved
- NVIDIA HGX H100 Preemptible Compute $1.99 per GPU per hour
- NVIDIA HGX H100 ON-Demand $3.99 per GPU per hour
- NVIDIA HGX H100 7-30 days $3.69 per GPU per hour
- NVIDIA HGX H100 31-90 days $3.45 per GPU per hour
- NVIDIA HGX H100 91-180 days $3.19 per GPU per hour
- On-demand hourly rates and reserved capacity
- Reservation terms shown as 7-30, 31-90, and 91-180 days, and 181+ days
Code Sandbox
- Per vCPU $0.0446 per hour
- Per GiB RAM $0.0149 per hour
- Customize a deployment of VM sandboxes for large development environments
Code Interpreter
- Session (60 minutes) $0.03 per session
- Execute LLM-generated code securely using the API
Managed Storage
- Shared Filesystem $0.16 GiB/month
- High-bandwidth, parallel filesystem colocated with your compute
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Runs commands — “await client.commands.run("npm install && npm run build")” source
- Choice of models — “Scale to 30 billion tokens per model with any serverless model or private deployment.” source
- API — “High-performance inference as APIs” source
- Runs models for you — “The fastest way to run open-source models on demand.” source
- Builds agents and workflows — “Build voice agents for production” source
- Traces and evaluates — “Measure model quality” source
Latest updates
- How to train your own Jev for $17
Launched the together/Tev1-4B-experimental classifier on Together’s serverless platform.
- Canary rollouts: upgrade models in production without downtime
Dedicated inference supports staged traffic ramps, metric gates, and automatic rollback for model upgrades.
- Together AI expands fine-tuning service with more models, live metrics, and finer controls
Together Fine-Tuning added more models, live experiment tracking, Expert LoRA, early stopping, dataset previews, and pre-flight validation.