tools
Groq - Fastest LLM Inference API | AI.info
Groq's LPU hardware delivers the fastest AI inference available — 800+ tokens/second. OpenAI-compatible API for Llama, DeepSeek, Mixtral, and more.

In inglese
GroqCloud lets developers and businesses call hosted AI models through APIs and a web console. It supports text generation, speech-to-text, text-to-speech, image-to-text, reasoning, structured outputs, prompt caching, and agentic tools.
Users can choose public, private, or co-cloud deployments, with service tiers for on-demand, flexible, and enterprise performance. A free tier is available; paid usage is token-based, while enterprise features such as custom models and dedicated support require a business plan.
Features
- Run hosted LLMs, speech, text-to-speech, and image-to-text models
- Use OpenAI-compatible chat completions and Responses APIs
- Add web search, website visiting, code execution, and browser automation
- Generate structured JSON outputs and use function calling
- Use prompt caching to reduce repeated-input costs and latency
- Process requests with on-demand, flex, batch, or performance tiers
- Connect external tools and services through MCP and built-in connectors
Use cases
- Build real-time chat and conversational applications
- Run voice transcription and text-to-speech workflows
- Create agents that search the web and execute code
- Add AI reasoning and structured extraction to business software
- Moderate text and images with hosted safety models
- Deploy latency-sensitive inference for production applications
Pros
Cons
Latest updates
- Added MCP Connectors (Beta)
Groq now supports Google Workspace connectors for Gmail, Google Calendar, and Google Drive using Model Context Protocol (MCP).
- Added OpenAI GPT-OSS-Safeguard 20B
This model helps classify text content based on customizable policies.
- Added Prompt Caching Enabled for GPT-OSS 120B
Automatic prompt caching is now live for openai/gpt-oss-120b; cache hits provide 50% cost savings on cached input tokens.
- Changed Python SDK v0.33.0, TypeScript SDK v0.34.0
Improved prompt caching support; added annotation/citation support to chat completion messages and streamed deltas.
Capabilities
- Choice of models — “Explore all available models on GroqCloud.” source
- API — “Get started with the Groq API” source
- Official SDKs — “import Groq from "groq-sdk";” source
- Runs models for you — “Hosted models are directly accessible through the GroqCloud Models API endpoint using the model IDs mentioned above.” source
Get it
Pricing
- Prices checked
- 2026-09-25