tools
Replicate
Replicate lets developers run, deploy, and scale machine learning models through a web interface or cloud API.

Replicate provides a web interface and API for running public machine learning models, including models for image, video, audio, and text tasks. Developers can integrate models into websites, chatbots, mobile apps, and other software.
Teams can also package and deploy custom models with Cog, create public or private models, and monitor predictions with logs and metrics. Usage is billed by compute time or model-specific input and output; private deployments may also incur setup and idle costs.
Features
- Run public machine learning models through a cloud API
- Run models from a browser-based web interface
- Search and compare public models
- Deploy custom models with the open-source Cog tool
- Create public or private custom models
- Automatically scale model deployments with demand
- Inspect prediction logs and performance metrics
Use cases
- Add image generation to a website or mobile app
- Build chatbots that call hosted language models
- Deploy a custom computer-vision model without managing GPUs
- Generate or transform video and audio through an API
- Test models in the browser before integrating them into software
Pros
Cons
Pricing
- Starting price
- $0.000025/sec
- Pricing checked
- 2026-09-19
CPU (Small)
$0.000025/sec; $0.09/hr
- 1x CPU
- 2GB RAM
CPU
$0.000100/sec; $0.36/hr
- 4x CPU
- 8GB RAM
Nvidia A100 (80GB) GPU
$0.001400/sec; $5.04/hr
- 1x GPU
- 10x CPU
- 80GB GPU RAM
- 144GB RAM
2x Nvidia A100 (80GB) GPU
$0.002800/sec; $10.08/hr
- 2x GPU
- 20x CPU
- 160GB GPU RAM
- 288GB RAM
Nvidia H100 GPU
$0.001525/sec; $5.49/hr
- 1x GPU
- 13x CPU
- 80GB GPU RAM
- 144GB RAM
Nvidia L40S GPU
$0.000975/sec; $3.51/hr
- 1x GPU
- 10x CPU
- 48GB GPU RAM
- 65GB RAM