Skip to content
AI.info

Tools

Cerebras Inference vs Fireworks AI

Cerebras Inference or Fireworks AI? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.

In short

  • Both: Choice of models, API, Runs models for you
  • Only Fireworks AI states: Command line, Official SDKs

Cerebras Inference

Cerebras Inference provides API access to fast cloud inference for language and multimodal models.

Plans

Developer Tier — OpenAI GPT OSS 120B

  • $0.35/M tokens
  • $0.75/M tokens
  • ~3000 tokens/s

Developer Tier — QWEN Qwen 3.8 27B

  • $0.99/M tokens
  • $1.49/M tokens
  • ~1,850 tokens/s

Prices checked 2026-09-25 on the maker’s page.

Capabilities

  • Choice of models — “choose the best model for your use case and requirements.” source
  • API — “OpenAI API compatibility lets developers build on Cerebras with just two code changes.” source
  • Runs models for you — “The Fastest AI Inference Cloud” source

Security

  • SOC 2 Type II — “SOC 2 Type 2” source
  • SOC 2 — “SOC 2 Report” source
  • GDPR — “LEGAL Data Processing Agreement” source

Latest updates

About Cerebras Inference

Fireworks AI

Fireworks AI hosts, fine-tunes, and serves open models through serverless and dedicated APIs.

Plans

Serverless Inference — Embeddings (up to 150M parameters)

  • $0.008 / 1M input tokens
  • Per-token pricing
  • Zero setup
  • No cold starts
  • High rate limits
  • Postpaid billing

Serverless Inference — Embeddings (150M–350M parameters)

  • $0.016 / 1M input tokens
  • Per-token pricing
  • Zero setup
  • No cold starts
  • High rate limits
  • Postpaid billing

Serverless Inference — Qwen3 8B

  • $0.1 / 1M input tokens
  • Per-token pricing
  • Zero setup
  • No cold starts
  • High rate limits
  • Postpaid billing

Managed Training — Models up to 16B parameters

  • $0.50 / 1M training tokens
  • $1.00 / 1M training tokens
  • $2.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Managed Training — Models 16.1B–80B

  • $3.00 / 1M training tokens
  • $6.00 / 1M training tokens
  • $12.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Managed Training — Models 80B–300B

  • $6.00 / 1M training tokens
  • $12.00 / 1M training tokens
  • $24.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Managed Training — Models over 300B

  • $10.00 / 1M training tokens
  • $20.00 / 1M training tokens
  • $40.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Serverless Training API — GLM 5.3

  • $4.86 / 1M Prefill
  • $0.972 / 1M Cached Prefill
  • $12.15 / 1M Sample
  • $14.58 / 1M Train
  • Shared, always-on trainer pool for LoRA training
  • No provisioning or idle cost
  • Pay only for tokens prefetched, sampled, and trained

Serverless Training API — Qwen 3.8 27B

  • $1.86 / 1M Prefill
  • $0.372 / 1M Cached Prefill
  • $5.595 / 1M Sample
  • $4.103 / 1M Train
  • Shared, always-on trainer pool for LoRA training
  • No provisioning or idle cost
  • Pay only for tokens prefetched, sampled, and trained

Serverless Training API — Kimi K3

  • $10.87 / 1M Prefill
  • $2.17 / 1M Cached Prefill
  • $27.11 / 1M Sample
  • $32.55 / 1M Train
  • Shared, always-on trainer pool for LoRA training
  • No provisioning or idle cost
  • Pay only for tokens prefetched, sampled, and trained

Prices checked 2026-09-25 on the maker’s page.

Capabilities

  • Command line — “Developers Model Library Docs CLI API Changelog” source
  • Choice of models — “Route to the best open or closed model for every task, and cut your AI coding spend 50 to 75%.” source
  • API — “Serverless. Pay per token with Priority and Fast options to meet your requirements. OpenAI and Anthropic compatible.” source
  • Official SDKs — “The Fireworks Training SDK lets us focus on our research instead of wrestling with infrastructure.” source
  • Runs models for you — “Serve the latest open models, or your own trained versions.” source

Security

  • SOC 2 Type II — “SOC 2 Type 2” source
  • SOC 2 — “SOC 2 Type 2” source
  • ISO 27001 — “ISO 27001 Certificate” source
  • ISO 42001 — “ISO 42001 Certificate” source
  • GDPR — “Compliance SOC 2 Type 2 HIPAA GDPR” source
  • HIPAA — “SOC 2 Type 2 HIPAA” source

Latest updates

About Fireworks AI