tools
NVIDIA NIM
NVIDIA NIM provides containerized, optimized inference microservices for deploying AI models on NVIDIA-accelerated infrastructure.

In inglese
NVIDIA NIM packages AI models, optimized inference engines, runtime dependencies, and standard APIs into deployable software containers. Teams can use hosted APIs for prototyping or self-host NIM in clouds, data centers, workstations, and edge environments.
Developers use NIM to build AI agents, virtual assistants, document-processing systems, shopping tools, and 3D configurators. Development access is free through the NVIDIA Developer Program; production use requires NVIDIA AI Enterprise licensing, and the license does not cover the underlying model.
Features
- Deploy prebuilt inference microservices with a single command
- Expose models through industry-standard APIs
- Run on NVIDIA-accelerated cloud, data center, workstation, and edge infrastructure
- Support models based on TensorRT-LLM, vLLM, and SGLang
- Run fine-tuned models with self-hosted NIM endpoints
- Scale deployments on Kubernetes and cloud service providers
- Use hosted APIs accelerated by DGX Cloud for prototyping
- Integrate NIM endpoints with OpenAI-compatible client code
Use cases
- Build AI virtual assistants for customer support and business processes
- Process documents with generative AI
- Create hyperpersonalized shopping experiences
- Deploy 3D product configurator applications
- Prototype AI agents with NVIDIA-hosted APIs
- Run production inference inside a controlled infrastructure environment
Pros
Cons
Capabilities
- Command line — “Deploy NIM for your model with a single command.” source
- Choice of models — “Deploy large language models (LLMs) supported by NVIDIA® TensorRT™-LLM, vLLM, or SGLang for low-latency, high-throughput inferencing on NVIDIA-accelerated infrastructure.” source
- Self-hosted — “NVIDIA NIM combines the ease of use and operational simplicity of managed APIs with the flexibility and security of self-hosting models on your preferred infrastructure.” source
- Runs models for you — “Get access to unlimited prototyping with hosted APIs for NIM accelerated by DGX Cloud, or download and self-host NIM microservices for research and development as part of the NVIDIA Developer program.” source
- Builds agents and workflows — “Build AI Agents With NIM” source
Get it
Pricing
- Starting price
- $1125/yr
- Prices checked
- 2026-09-24
Subscription (Includes support)
- $4,500 / GPU (1 year; List Pricing)
- $1,125 / GPU (1 year; EDU and Inception Pricing)
- $9,000 / GPU (2 years; List Pricing)
- $2,250 / GPU (2 years; EDU and Inception Pricing)
- $13,500 / GPU (3 years; List Pricing)
- $3,375 / GPU (3 years; EDU and Inception Pricing)
- $18,000 / GPU (4 years; List Pricing)
- $4,500 / GPU (4 years; EDU and Inception Pricing)
- $18,000 / GPU (5 years; List Pricing)
- $4,500 / GPU (5 years; EDU and Inception Pricing)
- Includes support
Perpetual
- $22,500 / GPU (5 years support; List Pricing)
- $5,625 / GPU (5 years support; EDU and Inception Pricing)
- 5 years support
Production
- $1 / hour / GPU + CSP Instance Cost(s)
- Consumption / Pay as you go
- Limited to 3 calls
Development
Free
- Free to use or BYOL + CSP Instance Cost(s)
- Developer Forum and Discord
Private Offer
Price on request
- 1- to 3-year subscription
- NVIDIA AI Enterprise Support