Vai al contenuto
AI.info

tools

Together AI - Open-Source Model API & Fine-Tuning Platform | AI.info

Together AI provides API access to 100+ open-source models including Llama 4 and DeepSeek. OpenAI-compatible with fine-tuning and dedicated endpoints.

Together AI - Open-Source Model API & Fine-Tuning Platform | AI.info

In inglese

Together AI is a developer platform for serverless and dedicated model inference, batch processing, fine-tuning, GPU clusters, sandboxes, and managed storage. It supports text, image, audio, video, transcription, embeddings, reranking, and moderation workloads.

Developers and AI teams use it to build applications, agents, media features, and production model deployments. Pricing is usage-based rather than subscription-based; dedicated infrastructure, compute, storage, and fine-tuning cost extra, and the platform requires at least $5 in purchased credits.

Features

  • Run serverless inference for text, image, audio, video, transcription, and embedding models
  • Process large workloads asynchronously with Batch Inference
  • Deploy models on dedicated, single-tenant GPU infrastructure
  • Fine-tune open models with supervised fine-tuning and preference optimization
  • Run GPU clusters for training and inference workloads
  • Create secure code sandboxes and code interpreter sessions through an API
  • Use managed storage with parallel filesystems and no egress fees

Use cases

  • Build production AI applications with model APIs
  • Deploy coding, research, and customer-support agents
  • Fine-tune open models for domain-specific behavior
  • Process batch text, image, audio, or video workloads
  • Run training and inference jobs on managed GPU clusters
  • Add voice, media generation, or multimodal features to products

Pros

    Cons

      Latest updates

      Capabilities

      • Runs commands — “await client.commands.run("npm install && npm run build")” source
      • Choice of models — “Scale to 30 billion tokens per model with any serverless model or private deployment.” source
      • API — “High-performance inference as APIs” source
      • Runs models for you — “The fastest way to run open-source models on demand.” source
      • Builds agents and workflows — “Build voice agents for production” source
      • Traces and evaluates — “Measure model quality” source

      Get it

      Pricing

      Starting price
      $0.16/mo
      Prices checked
      2026-09-25

      Serverless Inference

      • MiniMax M3 — $0.30 per 1M tokens (Input)
      • MiniMax M3 — $1.20 per 1M tokens (output)
      • Kimi K3 — $3.00 per 1M tokens (Input)
      • Kimi K3 — $15.00 per 1M tokens (output)
      • GLM-5.3-Flash — $0.15 per 1M tokens (Input)
      • GLM-5.3-Flash — $0.50 per 1M tokens (output)
      • GPT Image 2 — $0.053 per image
      • Wan 2.6 Image — $0.03 per image
      • ByteDance Seedance 2.5 — $0.115 per video
      • ByteDance Seedance 2.0 — $0.16 per video
      • NVIDIA Nemotron 3 ASR Streaming 0.6B — $0.0015 per audio minute
      • Whisper Large v3 — $0.0015 per audio minute
      • High-performance inference as APIs
      • Prices vary by model and task; the page also lists batch API prices.

      Dedicated Inference — NVIDIA HGX H100

      • $5.49 per gpu per hour
      • $3.99 per gpu per hour
      • Single-tenant GPU instances
      • Guaranteed performance (no sharing)
      • Support for custom models
      • Autoscaling & traffic spike handling

      Dedicated Inference — NVIDIA HGX B200

      • $8.99 per gpu per hour
      • Single-tenant GPU instances
      • Guaranteed performance (no sharing)
      • Support for custom models
      • Autoscaling & traffic spike handling

      Dedicated Inference — other hardware

      Price on request

      • NVIDIA HGX H200, NVIDIA HGX B300, NVIDIA GB200 NVL72, and NVIDIA GB300 NVL72
      • Contact sales

      GPU Clusters — On-demand

      • NVIDIA HGX B200 $8.19 per GPU per hour
      • NVIDIA HGX B300 $9.99 per GPU per hour
      • NVIDIA HGX H100 $3.99 per GPU per hour
      • NVIDIA HGX H200 $5.99 per GPU per hour
      • Pay-as-you-go GPU capacity on an hourly basis

      GPU Clusters — Preemptible and reserved

      • NVIDIA HGX H100 Preemptible Compute $1.99 per GPU per hour
      • NVIDIA HGX H100 ON-Demand $3.99 per GPU per hour
      • NVIDIA HGX H100 7-30 days $3.69 per GPU per hour
      • NVIDIA HGX H100 31-90 days $3.45 per GPU per hour
      • NVIDIA HGX H100 91-180 days $3.19 per GPU per hour
      • On-demand hourly rates and reserved capacity
      • Reservation terms shown as 7-30, 31-90, and 91-180 days, and 181+ days

      Code Sandbox

      • Per vCPU $0.0446 per hour
      • Per GiB RAM $0.0149 per hour
      • Customize a deployment of VM sandboxes for large development environments

      Code Interpreter

      • Session (60 minutes) $0.03 per session
      • Execute LLM-generated code securely using the API

      Managed Storage

      • Shared Filesystem $0.16 GiB/month
      • High-bandwidth, parallel filesystem colocated with your compute
      Official website