tools
Cerebras Inference
Cerebras Inference provides API access to fast cloud inference for language and multimodal models.

Cerebras Inference lets developers run supported AI models through an OpenAI-compatible API. It is intended for coding, research, voice, automation, agentic workflows, and other applications where response speed matters.
The service offers public endpoints, partner integrations, and dedicated enterprise endpoints. The product page advertises a free trial with $5 in credit; developer usage is pay per token, while enterprise access requires contacting Cerebras. Current pricing tier amounts were not rendered on the pricing page.
Features
- OpenAI API compatibility with two code changes
- Run supported language and multimodal models through cloud endpoints
- Public endpoints for models including GPT OSS 120B and Gemma 4 31B
- Dedicated enterprise endpoints for larger model families
- Multi-LoRA support for switching adapters per request
- Partner access through AWS, OpenRouter, Hugging Face, and Vercel
- Free trial with $5 in credit
Use cases
- Build responsive coding assistants and coding agents
- Run research and knowledge-work agents with lower response latency
- Create voice and real-time interactive applications
- Automate browser and computer-use workflows
- Serve specialized model behavior with multiple LoRA adapters