Skip to content
AI.info

tools

RunPod Serverless

RunPod Serverless deploys containerized AI inference endpoints that autoscale GPU workers and bill compute by the second.

RunPod Serverless

RunPod Serverless runs containerized inference workloads behind an API. It scales workers with request demand, can scale to zero when idle, and supports real-time and batch inference.

Developers and AI teams use it to serve image, speech, recommendation, and language models without managing GPU servers. Pricing is usage-based rather than subscription-based; GPU selection, storage, and deployment choices affect total cost.

Features

  • Autoscaling GPU endpoints that scale from zero to hundreds of workers
  • Per-second billing from worker start to full stop
  • FlashBoot cold-start optimization with sub-200ms cold starts on active endpoints
  • Batch inference for high-volume jobs
  • Up to 8 GPUs behind a single worker
  • Deploy with containers or the no-Docker Flash path
  • Network storage for model caching and persistent data
  • Webhooks, APIs, custom event triggers, logs, metrics, and tracing

Use cases

  • Serve image-generation models through an API
  • Run speech recognition or transcription workloads
  • Deploy language-model inference for applications
  • Process large image or document collections with batch inference
  • Scale AI agents and recommendation services with request demand

Pros

    Cons

      Pricing

      Starting price
      $0.58/hr
      Pricing checked
      2026-09-19

      B300

      $9.98/hr

      • 280 GB
      • Maximum throughput for big models

      B200

      $8.64/hr

      • 180 GB
      • Maximum throughput for big models

      H200

      $5.93/hr

      • 141 GB
      • Extreme throughput for big models

      RTX 6000 Pro

      $3.49/hr

      • 96 GB
      • High throughput for large model inference workloads

      H100

      $4.79/hr

      • 80 GB
      • Extreme throughput for big models

      A100

      $2.72/hr

      • 80 GB
      • High throughput GPU for inference
      Official website