Tools
Fireworks AI vs Modal
Fireworks AI or Modal? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In short
- Both: Choice of models, API, Official SDKs, Runs models for you
- Only Fireworks AI states: Command line
- Only Modal states: Agent that edits files, Builds agents and workflows, Traces and evaluates
Fireworks AI
Fireworks AI hosts, fine-tunes, and serves open models through serverless and dedicated APIs.
Plans
Serverless Inference — Embeddings (up to 150M parameters)
- $0.008 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Embeddings (150M–350M parameters)
- $0.016 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Qwen3 8B
- $0.1 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Managed Training — Models up to 16B parameters
- $0.50 / 1M training tokens
- $1.00 / 1M training tokens
- $2.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 16.1B–80B
- $3.00 / 1M training tokens
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 80B–300B
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- $24.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models over 300B
- $10.00 / 1M training tokens
- $20.00 / 1M training tokens
- $40.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Serverless Training API — GLM 5.3
- $4.86 / 1M Prefill
- $0.972 / 1M Cached Prefill
- $12.15 / 1M Sample
- $14.58 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Qwen 3.8 27B
- $1.86 / 1M Prefill
- $0.372 / 1M Cached Prefill
- $5.595 / 1M Sample
- $4.103 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Kimi K3
- $10.87 / 1M Prefill
- $2.17 / 1M Cached Prefill
- $27.11 / 1M Sample
- $32.55 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Command line — “Developers Model Library Docs CLI API Changelog” source
- Choice of models — “Route to the best open or closed model for every task, and cut your AI coding spend 50 to 75%.” source
- API — “Serverless. Pay per token with Priority and Fast options to meet your requirements. OpenAI and Anthropic compatible.” source
- Official SDKs — “The Fireworks Training SDK lets us focus on our research instead of wrestling with infrastructure.” source
- Runs models for you — “Serve the latest open models, or your own trained versions.” source
Security
- SOC 2 Type II — “SOC 2 Type 2” source
- SOC 2 — “SOC 2 Type 2” source
- ISO 27001 — “ISO 27001 Certificate” source
- ISO 42001 — “ISO 42001 Certificate” source
- GDPR — “Compliance SOC 2 Type 2 HIPAA GDPR” source
- HIPAA — “SOC 2 Type 2 HIPAA” source
Latest updates
- Serverless pricing update: DeepSeek V4.1 Flash
Serverless pricing for DeepSeek V4.1 Flash changes; dedicated deployment and Reserved Throughput pricing is unaffected.
- New deployment creation flags: deploymentShape: "default" and acceptShapelessRisk
Create Deployment adds deploymentShape: "default" to pick a validated deployment shape and acceptShapelessRisk to create without a shape.
- Upcoming Serverless deprecation: older DeepSeek, GLM, Muse, and Kimi models
Several older Serverless models will be decommissioned on September 25, 2026; migrate to a recommended replacement before then.
Modal
Modal is a serverless cloud platform for running AI inference, training, batch jobs, and isolated code sandboxes.
Plans
Starter
Free
- $30 / month free compute
- 3 workspace seats included
- 100 containers + 10 GPU concurrency
- Scheduled and Web Functions (limited)
- Real-time metrics and logs
- Region selection
Team
- $250 + compute / month
- $100 / month free compute
- Unlimited seats
- 5000 containers + 50 GPU concurrency
- Unlimited Scheduled Functions
- Custom domains
- Static IP proxy
Enterprise
Price on request
- Volume-based discounts
- Unlimited seats
- Higher GPU concurrency
- Embedded ML engineering services
- Environment-level budgets
- Support via private Slack
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Agent that edits files — “Autonomous agents with the right tools, context, and credentials already in place — running securely in a full, isolated dev environment.” source
- Choice of models — “Run any model or inference engine on H100s, A100s, A10Gs and more.” source
- API — “Serve your own LLM API” source
- Official SDKs — “Modal SDK” source
- Runs models for you — “Deploy and scale inference for LLMs, audio, image/video generation.” source
- Builds agents and workflows — “Designed to scale agents.” source
- Traces and evaluates — “Debug fast by zooming into metrics, logs, and live statuses of specific inference calls.” source
Latest updates
- Network egress usage broken down by app
The Usage page now shows network egress per app.
- New controls for unauthenticated web endpoints
Filter by authentication, audit unauthenticated deployments, and block them per environment.
- User groups for RBAC-enabled workspaces
Assign environment roles to groups of members, created manually or synced from your identity provider via SCIM.