tools
Replicate
Replicate lets developers run, deploy, and scale machine learning models through a web interface or cloud API.

In inglese
Replicate provides a web interface and API for running public machine learning models, including models for image, video, audio, and text tasks. Developers can integrate models into websites, chatbots, mobile apps, and other software.
Teams can also package and deploy custom models with Cog, create public or private models, and monitor predictions with logs and metrics. Usage is billed by compute time or model-specific input and output; private deployments may also incur setup and idle costs.
Features
- Run public machine learning models through a cloud API
- Run models from a browser-based web interface
- Search and compare public models
- Deploy custom models with the open-source Cog tool
- Create public or private custom models
- Automatically scale model deployments with demand
- Inspect prediction logs and performance metrics
Use cases
- Add image generation to a website or mobile app
- Build chatbots that call hosted language models
- Deploy a custom computer-vision model without managing GPUs
- Generate or transform video and audio through an API
- Test models in the browser before integrating them into software
Pros
Cons
Latest updates
- Agent skills for Replicate
Replicate now publishes agent skills, markdown instruction files that give coding assistants knowledge about working with AI models on Replicate.
- Fallback model for Nano Banana Pro
Nano Banana Pro can fall back to Seedream 5.0 lite when Google’s API is at capacity.
- MCP server auto-discovery
Replicate’s MCP server can now be discovered automatically through the official MCP Registry.
- Filter predictions by source
You can now filter the list predictions API endpoint to show only predictions created through the web interface.
- The little things, week ending December 19, 2025
Improved the reliability of google/nano-banana and google/nano-banana-pro; added automatic llms.txt generation for documentation.
Capabilities
- Choice of models — “Compare models in the Playground” source
- API — “Replicate - Run AI with an API” source
- Official SDKs — “import Replicate from "replicate"” source
- Runs models for you — “Cog takes care of generating an API server and deploying it on a big cluster in the cloud.” source
- Traces and evaluates — “Metrics let you keep an eye on how your models are performing, and logs let you zoom in on particular predictions to debug how your model is behaving.” source
Get it
Pricing
- Prices checked
- 2026-09-25
black-forest-labs / flux-1.1-pro
- $0.04 / output image
- Text-to-image model
- Excellent image quality, prompt adherence, and output diversity
black-forest-labs / flux-dev
- $0.025 / output image
- 12 billion parameter rectified flow transformer
- Generates images from text descriptions
black-forest-labs / flux-schnell
- $3.00 / thousand output images
- Fast image generation
- Tailored for local development and personal use
deepseek-ai / deepseek-r1
- $0.01 / thousand output tokens
- $3.75 / million input tokens
- Reasoning model trained with reinforcement learning
- On par with OpenAI o1
ideogram-ai / ideogram-v3-quality
- $0.09 / output image
- Highest quality Ideogram v3 model
- Creates images with realism, creative designs, and consistent styles
recraft-ai / recraft-v3
- $0.04 / output image
- Text-to-image model
- Generates long texts and images in a wide list of styles
wavespeedai / wan-2.1-i2v-480p
- $0.09 / second of output video
- Accelerated inference for Wan 2.1 14B image to video
wavespeedai / wan-2.1-i2v-720p
- $0.25 / second of output video
- Accelerated inference for Wan 2.1 14B image to video with high resolution
CPU (Small)
- $ 0.000025 /sec
- $ 0.09 /hr
- cpu-small
- CPU 1x
- RAM 2GB
CPU
- $ 0.36 /hr
- cpu
- CPU 4x
- RAM 8GB