tools
RunPod Serverless
RunPod Serverless deploys containerized AI inference endpoints that autoscale GPU workers and bill compute by the second.

In inglese
RunPod Serverless runs containerized inference workloads behind an API. It scales workers with request demand, can scale to zero when idle, and supports real-time and batch inference.
Developers and AI teams use it to serve image, speech, recommendation, and language models without managing GPU servers. Pricing is usage-based rather than subscription-based; GPU selection, storage, and deployment choices affect total cost.
Features
- Autoscaling GPU endpoints that scale from zero to hundreds of workers
- Per-second billing from worker start to full stop
- FlashBoot cold-start optimization with sub-200ms cold starts on active endpoints
- Batch inference for high-volume jobs
- Up to 8 GPUs behind a single worker
- Deploy with containers or the no-Docker Flash path
- Network storage for model caching and persistent data
- Webhooks, APIs, custom event triggers, logs, metrics, and tracing
Use cases
- Serve image-generation models through an API
- Run speech recognition or transcription workloads
- Deploy language-model inference for applications
- Process large image or document collections with batch inference
- Scale AI agents and recommendation services with request demand
Pros
Cons
Latest updates
- Global Volumes - Beta
Create a global volume once and mount it at startup across any region, with no copying data between data centers required.
- Nano Banana Edit
The Nano Banana Edit public endpoint will be retired on September 28, 2026. Migrate to Nano Banana 2 Edit.
- Sales tax and tax ID support
Runpod now collects sales tax on credit purchases in applicable jurisdictions. Add a business tax ID at checkout or in account settings.
- Batch Jobs (BETA)
Submit large sets of inference requests to a Serverless endpoint as a single managed unit. Create a batch, finalize it to start processing, and poll for status and…
- REST API v2
REST API v2 is now generally available. It reorganizes resource paths, standardizes request and response shapes, and adds catalog endpoints, pod log streaming, and…
Capabilities
- Command line — “Deploy and manage directly from your terminal.” source
- Choice of models — “Instant access to pre-deployed AI models via API. No infrastructure setup required.” source
- API — “Serverless GPU endpoints run containerized inference workloads behind an API and scale workers based on demand.” source
- Official SDKs — “CLI & SDKs.” source
- Runs models for you — “Deploy any containerized model as an autoscaling GPU endpoint without managing servers.” source
- Builds agents and workflows — “Build intelligent agent-based systems and workflows.” source
- Traces and evaluates — “Runpod offers a comprehensive monitoring dashboard with real-time logging and distributed tracing for your serverless functions.” source
Get it
Pricing
- Prices checked
- 2026-09-25
B300
- $ 9.98 /hr
- Maximum throughput for big models.
B200
- $ 8.64 /hr
- Maximum throughput for big models.
H200
- $ 5.93 /hr
- Extreme throughput for big models.
RTX 6000 Pro
- $ 3.49 /hr
- High throughput for large model inference workloads.
H100
- $ 4.79 /hr
- Extreme throughput for big models.
A100
- $ 2.72 /hr
- High throughput GPU, yet still very cost-effective.
L40, L40S, 6000 Ada, MIG 48GB
- $ 1.75 /hr
- Extreme inference throughput on LLMs like Llama 3 7B.
A6000, A40
- $ 1.22 /hr
- A cost-effective option for running big models.
5090
- $ 1.58 /hr
- Extreme throughput for small-to-medium models.
RTX PRO 4500 Blackwell
- $ 1.15 /hr
- Cost-effective Blackwell inference for 32GB workloads.