tools
Cerebras Inference
Cerebras Inference provides API access to fast cloud inference for language and multimodal models.

In inglese
Cerebras Inference lets developers run supported AI models through an OpenAI-compatible API. It is intended for coding, research, voice, automation, agentic workflows, and other applications where response speed matters.
The service offers public endpoints, partner integrations, and dedicated enterprise endpoints. The product page advertises a free trial with $5 in credit; developer usage is pay per token, while enterprise access requires contacting Cerebras. Current pricing tier amounts were not rendered on the pricing page.
Features
- OpenAI API compatibility with two code changes
- Run supported language and multimodal models through cloud endpoints
- Public endpoints for models including GPT OSS 120B and Gemma 4 31B
- Dedicated enterprise endpoints for larger model families
- Multi-LoRA support for switching adapters per request
- Partner access through AWS, OpenRouter, Hugging Face, and Vercel
- Free trial with $5 in credit
Use cases
- Build responsive coding assistants and coding agents
- Run research and knowledge-work agents with lower response latency
- Create voice and real-time interactive applications
- Automate browser and computer-use workflows
- Serve specialized model behavior with multiple LoRA adapters
Pros
Cons
Latest updates
- Temporary increase to Qwen 3.8 27B total token rate limit
The Developer tier total token rate limit for qwen-3.8-27b is temporarily increased from 450K to 750K tokens per minute.
- Gemma 4 31B availability changes on public endpoints
gemma-4-31b is no longer available on Cerebras public endpoints. Gemma 4 31B remains available on Dedicated Endpoints.
- Qwen 3.8 27B available on public endpoints
qwen-3.8-27b is now available on Cerebras public endpoints.
- Flowise integration moved to legacy status
We removed Flowise from the supported integrations directory and navigation. The legacy Flowise guide remains available for existing self-hosted deployments.
- GLM 4.7 reasoning_logprobs default will not ship before deprecation
zai-glm-4.7 will keep its existing default logprobs response shape until deprecation. Requests that explicitly set the X-Cerebras-Version-Patch: 2 header still receive…
Capabilities
Security
Pricing
- Prices checked
- 2026-09-25
Developer Tier — OpenAI GPT OSS 120B
- $0.35/M tokens
- $0.75/M tokens
- ~3000 tokens/s
Developer Tier — QWEN Qwen 3.8 27B
- $0.99/M tokens
- $1.49/M tokens
- ~1,850 tokens/s