tools
Ollama — Run Any LLM Locally for Free | 2026 Review
Ollama lets you run Llama 3.3, Mistral, Phi-4, DeepSeek and 100+ models locally with one command. OpenAI-compatible API, 100% private, free forever.

In inglese
Ollama lets developers download, run, and manage open models on macOS, Windows, and Linux. It includes a command-line interface, REST API, model library, desktop app, and integrations with coding tools such as Claude Code, Codex, OpenCode, and OpenClaw.
Local models run on your own hardware without usage charges. Cloud models require an account and use monthly plan credits or pay-as-you-go token pricing; larger teams can buy Enterprise plans with custom pricing.
Features
- Run open models locally on macOS, Windows, and Linux
- Use cloud models with monthly usage credits and token pricing
- Connect coding agents with the ollama launch command
- Expose a REST API for running and managing models
- Use Python and JavaScript libraries
- Support tool calling and structured outputs
- Import models and define them with Modelfiles
- MIT-licensed open-source codebase
Use cases
- Run private AI assistants without sending local prompts to a provider
- Build applications with the Ollama REST API
- Connect open models to coding agents and developer tools
- Prototype chat, retrieval, and automation workflows locally
- Serve models on GPUs or CPUs across supported operating systems
Pros
Cons
Latest updates
- v0.34.4 (v0.34.4)
Structured outputs on thinking models now apply in one pass; fixed model lookup and macOS app issues and changed Apple Silicon processing.
- v0.34.3 (v0.34.3)
GET /api/show now exposes model thinking controls and defaults; Nemotron H vision models are supported on Apple Silicon with MLX.
- v0.34.2 (v0.34.2)
Added first-run sign-in or local setup and an ollama://apps link to open the desktop app’s Apps page on macOS and Windows.
Capabilities
- VS Code — “Claude Code Codex OpenCode Hermes Agent OpenClaw VS Code Pi n8n” source
- Choice of models — “Switch models without changing your workflow.” source
- Self-hosted — “Run models locally” source
- No code retained — “Every request runs on dedicated compute in the US and Europe, plus Singapore for a limited set of Qwen models, with zero data retention.” source
- API — “Works with popular coding agents, including Claude Code and Codex, plus an API for your own tools” source
- Runs models for you — “Ollama hosts models and compute resources primarily in the United States.” source
Get it
Pricing
- Starting price
- $16.67/mo
- Prices checked
- 2026-09-26
Free
Free
- Run models locally
- Starter usage credits included
- Includes access to starter models
- Add credits to unlock all models
- No service fees
Pro
- $20 / mo.
- $200/yr billed annually
- $16.67 / mo. billed annually
- $60 of usage credits per month
- Access to larger pro models
- Run multiple models concurrently
- Fast mode (coming soon)
Max
- $100 / mo.
- $300 of usage credits per month
- Early access to the newest models
- 10 concurrent requests
- Everything in Pro
Team
- $500 / mo.
- Unlimited users
- $1,000 of usage credits per month, shared across the team
- Centralized billing and administration
- Priority support
- Shared projects, skills, and instructions (coming soon)
Enterprise
Price on request
- Model access controls
- Limit team access to specific models
- Set cost budgets for users and API keys
- Private Slack channel with dedicated support
- Custom security questionnaires