Vai al contenuto
AI.info

tools

Cerebras Inference

Cerebras Inference provides API access to fast cloud inference for language and multimodal models.

Cerebras Inference

In inglese

Cerebras Inference lets developers run supported AI models through an OpenAI-compatible API. It is intended for coding, research, voice, automation, agentic workflows, and other applications where response speed matters.

The service offers public endpoints, partner integrations, and dedicated enterprise endpoints. The product page advertises a free trial with $5 in credit; developer usage is pay per token, while enterprise access requires contacting Cerebras. Current pricing tier amounts were not rendered on the pricing page.

Features

  • OpenAI API compatibility with two code changes
  • Run supported language and multimodal models through cloud endpoints
  • Public endpoints for models including GPT OSS 120B and Gemma 4 31B
  • Dedicated enterprise endpoints for larger model families
  • Multi-LoRA support for switching adapters per request
  • Partner access through AWS, OpenRouter, Hugging Face, and Vercel
  • Free trial with $5 in credit

Use cases

  • Build responsive coding assistants and coding agents
  • Run research and knowledge-work agents with lower response latency
  • Create voice and real-time interactive applications
  • Automate browser and computer-use workflows
  • Serve specialized model behavior with multiple LoRA adapters

Pros

    Cons

      Latest updates

      Capabilities

      • Choice of models — “choose the best model for your use case and requirements.” source
      • API — “OpenAI API compatibility lets developers build on Cerebras with just two code changes.” source
      • Runs models for you — “The Fastest AI Inference Cloud” source

      Security

      • SOC 2 Type II — “SOC 2 Type 2” source
      • SOC 2 — “SOC 2 Report” source
      • GDPR — “LEGAL Data Processing Agreement” source

      Pricing

      Prices checked
      2026-09-25

      Developer Tier — OpenAI GPT OSS 120B

      • $0.35/M tokens
      • $0.75/M tokens
      • ~3000 tokens/s

      Developer Tier — QWEN Qwen 3.8 27B

      • $0.99/M tokens
      • $1.49/M tokens
      • ~1,850 tokens/s
      Official website