Skip to content
AI.info

tools

SambaNova Cloud

SambaCloud is a hosted API platform for running open-source AI models, with OpenAI-compatible endpoints, model integrations, and token-based pricing.

SambaNova Cloud

SambaCloud provides hosted inference for models including DeepSeek, Llama, Qwen, Gemma, MiniMax, and gpt-oss. Developers can access the models through an OpenAI-compatible API, a web playground, integrations, and starter kits.

Plans include Free, Developer, and Enterprise options. Usage is billed by model and token volume; Enterprise customers can request subscription pricing, custom rate limits, on-request models, and features such as BYOC. It is an inference service, not a tool for training or hosting arbitrary model code.

Features

  • OpenAI-compatible API endpoints
  • Hosted inference for open-source models
  • Web playground for testing models
  • Support for text, image, and audio-capable models
  • Integrations with CrewAI, Hugging Face, Cline, and AWS
  • Automatic prompt caching for MiniMax-M2.7
  • Production and preview model access
  • Enterprise options including BYOC and custom rate limits

Use cases

  • Build applications with hosted open-source language models
  • Prototype prompts and model calls in the web playground
  • Add model inference to coding and agent workflows
  • Run support bots using repeated documentation prompts
  • Create applications that reuse long prompts with prompt caching

Pros

    Cons

      Pricing

      Starting price
      Free
      Pricing checked
      2026-09-19

      MiniMax-M2.7

      0.60 USD input / 2.40 USD output per 1M tokens

      • Cached input: 0.06 USD
      • Input: 0.60 USD per 1M tokens
      • Output: 2.40 USD per 1M tokens

      DeepSeek-V3.1

      3 USD input / 4.50 USD output per 1M tokens

      • Input: 3 USD per 1M tokens
      • Output: 4.50 USD per 1M tokens

      DeepSeek-V3.2

      3 USD input / 4.50 USD output per 1M tokens

      • Input: 3 USD per 1M tokens
      • Output: 4.50 USD per 1M tokens

      gemma-4-31B-it

      0.38 USD input / 1.15 USD output per 1M tokens

      • Input: 0.38 USD per 1M tokens
      • Output: 1.15 USD per 1M tokens

      gpt-oss-120b

      0.22 USD input / 0.59 USD output per 1M tokens

      • Input: 0.22 USD per 1M tokens
      • Output: 0.59 USD per 1M tokens

      Meta-Llama-3.3-70B-Instruct

      0.60 USD input / 1.20 USD output per 1M tokens

      • Input: 0.60 USD per 1M tokens
      • Output: 1.20 USD per 1M tokens
      Official website