Vai al contenuto
AI.info

Tools

Fireworks AI vs Unsloth

Fireworks AI or Unsloth? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.

In inglese

In short

  • Both: Command line, Choice of models, API
  • Only Fireworks AI states: Official SDKs, Runs models for you
  • Only Unsloth states: Chat about your code, Runs commands, Self-hosted, Search over your data

Fireworks AI

Fireworks AI hosts, fine-tunes, and serves open models through serverless and dedicated APIs.

Plans

Serverless Inference — Embeddings (up to 150M parameters)

  • $0.008 / 1M input tokens
  • Per-token pricing
  • Zero setup
  • No cold starts
  • High rate limits
  • Postpaid billing

Serverless Inference — Embeddings (150M–350M parameters)

  • $0.016 / 1M input tokens
  • Per-token pricing
  • Zero setup
  • No cold starts
  • High rate limits
  • Postpaid billing

Serverless Inference — Qwen3 8B

  • $0.1 / 1M input tokens
  • Per-token pricing
  • Zero setup
  • No cold starts
  • High rate limits
  • Postpaid billing

Managed Training — Models up to 16B parameters

  • $0.50 / 1M training tokens
  • $1.00 / 1M training tokens
  • $2.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Managed Training — Models 16.1B–80B

  • $3.00 / 1M training tokens
  • $6.00 / 1M training tokens
  • $12.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Managed Training — Models 80B–300B

  • $6.00 / 1M training tokens
  • $12.00 / 1M training tokens
  • $24.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Managed Training — Models over 300B

  • $10.00 / 1M training tokens
  • $20.00 / 1M training tokens
  • $40.00 / 1M training tokens
  • Supervised and preference fine-tuning
  • Serve fine-tuned models for the same price as base models

Serverless Training API — GLM 5.3

  • $4.86 / 1M Prefill
  • $0.972 / 1M Cached Prefill
  • $12.15 / 1M Sample
  • $14.58 / 1M Train
  • Shared, always-on trainer pool for LoRA training
  • No provisioning or idle cost
  • Pay only for tokens prefetched, sampled, and trained

Serverless Training API — Qwen 3.8 27B

  • $1.86 / 1M Prefill
  • $0.372 / 1M Cached Prefill
  • $5.595 / 1M Sample
  • $4.103 / 1M Train
  • Shared, always-on trainer pool for LoRA training
  • No provisioning or idle cost
  • Pay only for tokens prefetched, sampled, and trained

Serverless Training API — Kimi K3

  • $10.87 / 1M Prefill
  • $2.17 / 1M Cached Prefill
  • $27.11 / 1M Sample
  • $32.55 / 1M Train
  • Shared, always-on trainer pool for LoRA training
  • No provisioning or idle cost
  • Pay only for tokens prefetched, sampled, and trained

Prices checked 2026-09-25 on the maker’s page.

Capabilities

  • Command line — “Developers Model Library Docs CLI API Changelog” source
  • Choice of models — “Route to the best open or closed model for every task, and cut your AI coding spend 50 to 75%.” source
  • API — “Serverless. Pay per token with Priority and Fast options to meet your requirements. OpenAI and Anthropic compatible.” source
  • Official SDKs — “The Fireworks Training SDK lets us focus on our research instead of wrestling with infrastructure.” source
  • Runs models for you — “Serve the latest open models, or your own trained versions.” source

Security

  • SOC 2 Type II — “SOC 2 Type 2” source
  • SOC 2 — “SOC 2 Type 2” source
  • ISO 27001 — “ISO 27001 Certificate” source
  • ISO 42001 — “ISO 42001 Certificate” source
  • GDPR — “Compliance SOC 2 Type 2 HIPAA GDPR” source
  • HIPAA — “SOC 2 Type 2 HIPAA” source

Latest updates

About Fireworks AI

Unsloth

Unsloth runs and fine-tunes AI models locally through a free, open-source desktop app, web UI, and code-based tools.

Plans

Free

Free

  • Open-source
  • Supports Mistral and Gemma
  • Supports Llama 1, 2, and 3
  • Supports 4 bit and 16 bit LoRA

unsloth Pro

Contact us

  • 2.5x number of GPUs faster than FA2
  • 20% less memory than OSS
  • Enhanced MultiGPU support
  • Up to 8 GPUs support

unsloth Enterprise

Contact us

  • 32x number of GPUs faster than FA2
  • Up to 30% accuracy
  • 5x faster inference
  • Full training
  • Multi-node support
  • Customer support

Prices checked 2026-09-24 on the maker’s page.

Capabilities

  • Chat about your code — “Download a model and start chatting in minutes.” source
  • Runs commands — “Execute Bash and Python in a secure sandbox so models can run code, test results and complete real tasks locally.” source
  • Command line — “then run unsloth start claude .” source
  • Choice of models — “Discover, manage and download the right quantization for your device from the built-in model hub.” source
  • Self-hosted — “Open-source. Free. 100% Local.” source
  • API — “Unsloth also exposes an OpenAI-compatible API, so existing apps, scripts and SDKs can connect to your local models through a familiar interface.” source
  • Search over your data — “Private web search, deep research, RAG, MCP + exports (NVFP4, GGUF)” source

Latest updates

  • Qwen-Image-2.1 + Skills (Qwen-Image-2.1)

    Adds local Qwen-Image-2.1, custom Agent Skills, draggable chats, faster reasoning blocks, and Linux update and installation options.

  • Qwen-Image-2.1 + Skills (Qwen-Image-2.1)

    Adds local Qwen-Image-2.1, custom Agent Skills, draggable chats, faster reasoning blocks, and Linux update and installation options.

  • Docker + Multi User + AMD Support (v0.1.810-beta)

    Adds a Docker image with NVIDIA and AMD support, multi-user accounts, diffusion support, ARM64 Windows CUDA support, and RDNA1/2 support.

About Unsloth