tools
Fireworks AI Review 2026 — Fastest Open-Source LLM Inference API
Fireworks AI delivers the fastest inference for Llama, DeepSeek, and 300+ open-source models. Sub-200ms latency, OpenAI-compatible API, function calling. Free tier available.

In inglese
Fireworks AI provides infrastructure for running open models and custom-trained models. Developers can use serverless inference, dedicated GPU deployments, fine-tuning, embeddings, tool calling, and multimodal models through APIs and a developer toolkit.
It is used by teams building coding assistants, conversational systems, agents, search, enterprise RAG, and multimodal applications. Usage is metered by tokens, training tokens, or GPU time; enterprise deployments and some dedicated services require contacting Fireworks.
Features
- Run open models with serverless inference and no GPU setup
- Deploy models on dedicated GPUs with per-second billing
- Fine-tune models with LoRA, DPO, SFT, and reinforcement fine-tuning
- Serve hundreds of fine-tuned variants using Multi-LoRA
- Build agents with JSON mode, grammar mode, and tool calling
- Use text, vision, audio, speech, image, and video models
- Access OpenAI- and Anthropic-compatible inference APIs
Use cases
- Build coding assistants and code-generation products
- Run conversational AI and customer-support systems
- Create agents that call APIs and retain context
- Add semantic search and enterprise RAG to applications
- Fine-tune models on proprietary business data
- Process images, documents, audio, and video
Pros
Cons
Latest updates
- Serverless pricing update: DeepSeek V4.1 Flash
Serverless pricing for DeepSeek V4.1 Flash changes; dedicated deployment and Reserved Throughput pricing is unaffected.
- New deployment creation flags: deploymentShape: "default" and acceptShapelessRisk
Create Deployment adds deploymentShape: "default" to pick a validated deployment shape and acceptShapelessRisk to create without a shape.
- Upcoming Serverless deprecation: older DeepSeek, GLM, Muse, and Kimi models
Several older Serverless models will be decommissioned on September 25, 2026; migrate to a recommended replacement before then.
- Training cost estimator
The training cost estimator helps estimate what a training job will cost before you run it.
- Training skill for coding agents
A Fireworks training skill is available for coding agents to plan a run, estimate cost, and wait for approval before spend.
Capabilities
- Command line — “Developers Model Library Docs CLI API Changelog” source
- Choice of models — “Route to the best open or closed model for every task, and cut your AI coding spend 50 to 75%.” source
- API — “Serverless. Pay per token with Priority and Fast options to meet your requirements. OpenAI and Anthropic compatible.” source
- Official SDKs — “The Fireworks Training SDK lets us focus on our research instead of wrestling with infrastructure.” source
- Runs models for you — “Serve the latest open models, or your own trained versions.” source
Get it
Security
Pricing
- Prices checked
- 2026-09-25
Serverless Inference — Embeddings (up to 150M parameters)
- $0.008 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Embeddings (150M–350M parameters)
- $0.016 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Serverless Inference — Qwen3 8B
- $0.1 / 1M input tokens
- Per-token pricing
- Zero setup
- No cold starts
- High rate limits
- Postpaid billing
Managed Training — Models up to 16B parameters
- $0.50 / 1M training tokens
- $1.00 / 1M training tokens
- $2.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 16.1B–80B
- $3.00 / 1M training tokens
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models 80B–300B
- $6.00 / 1M training tokens
- $12.00 / 1M training tokens
- $24.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Managed Training — Models over 300B
- $10.00 / 1M training tokens
- $20.00 / 1M training tokens
- $40.00 / 1M training tokens
- Supervised and preference fine-tuning
- Serve fine-tuned models for the same price as base models
Serverless Training API — GLM 5.3
- $4.86 / 1M Prefill
- $0.972 / 1M Cached Prefill
- $12.15 / 1M Sample
- $14.58 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Qwen 3.8 27B
- $1.86 / 1M Prefill
- $0.372 / 1M Cached Prefill
- $5.595 / 1M Sample
- $4.103 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained
Serverless Training API — Kimi K3
- $10.87 / 1M Prefill
- $2.17 / 1M Cached Prefill
- $27.11 / 1M Sample
- $32.55 / 1M Train
- Shared, always-on trainer pool for LoRA training
- No provisioning or idle cost
- Pay only for tokens prefetched, sampled, and trained