tools
SambaNova Cloud
SambaCloud is a hosted API platform for running open-source AI models, with OpenAI-compatible endpoints, model integrations, and token-based pricing.

SambaCloud provides hosted inference for models including DeepSeek, Llama, Qwen, Gemma, MiniMax, and gpt-oss. Developers can access the models through an OpenAI-compatible API, a web playground, integrations, and starter kits.
Plans include Free, Developer, and Enterprise options. Usage is billed by model and token volume; Enterprise customers can request subscription pricing, custom rate limits, on-request models, and features such as BYOC. It is an inference service, not a tool for training or hosting arbitrary model code.
Features
- OpenAI-compatible API endpoints
- Hosted inference for open-source models
- Web playground for testing models
- Support for text, image, and audio-capable models
- Integrations with CrewAI, Hugging Face, Cline, and AWS
- Automatic prompt caching for MiniMax-M2.7
- Production and preview model access
- Enterprise options including BYOC and custom rate limits
Use cases
- Build applications with hosted open-source language models
- Prototype prompts and model calls in the web playground
- Add model inference to coding and agent workflows
- Run support bots using repeated documentation prompts
- Create applications that reuse long prompts with prompt caching
Pros
Cons
Pricing
- Starting price
- Free
- Pricing checked
- 2026-09-19
MiniMax-M2.7
0.60 USD input / 2.40 USD output per 1M tokens
- Cached input: 0.06 USD
- Input: 0.60 USD per 1M tokens
- Output: 2.40 USD per 1M tokens
DeepSeek-V3.1
3 USD input / 4.50 USD output per 1M tokens
- Input: 3 USD per 1M tokens
- Output: 4.50 USD per 1M tokens
DeepSeek-V3.2
3 USD input / 4.50 USD output per 1M tokens
- Input: 3 USD per 1M tokens
- Output: 4.50 USD per 1M tokens
gemma-4-31B-it
0.38 USD input / 1.15 USD output per 1M tokens
- Input: 0.38 USD per 1M tokens
- Output: 1.15 USD per 1M tokens
gpt-oss-120b
0.22 USD input / 0.59 USD output per 1M tokens
- Input: 0.22 USD per 1M tokens
- Output: 0.59 USD per 1M tokens
Meta-Llama-3.3-70B-Instruct
0.60 USD input / 1.20 USD output per 1M tokens
- Input: 0.60 USD per 1M tokens
- Output: 1.20 USD per 1M tokens