tools
Groq - Fastest LLM Inference API | AI.info
Groq's LPU hardware delivers the fastest AI inference available — 800+ tokens/second. OpenAI-compatible API for Llama, DeepSeek, Mixtral, and more.

GroqCloud lets developers and businesses call hosted AI models through APIs and a web console. It supports text generation, speech-to-text, text-to-speech, image-to-text, reasoning, structured outputs, prompt caching, and agentic tools.
Users can choose public, private, or co-cloud deployments, with service tiers for on-demand, flexible, and enterprise performance. A free tier is available; paid usage is token-based, while enterprise features such as custom models and dedicated support require a business plan.
Features
- Run hosted LLMs, speech, text-to-speech, and image-to-text models
- Use OpenAI-compatible chat completions and Responses APIs
- Add web search, website visiting, code execution, and browser automation
- Generate structured JSON outputs and use function calling
- Use prompt caching to reduce repeated-input costs and latency
- Process requests with on-demand, flex, batch, or performance tiers
- Connect external tools and services through MCP and built-in connectors
Use cases
- Build real-time chat and conversational applications
- Run voice transcription and text-to-speech workflows
- Create agents that search the web and execute code
- Add AI reasoning and structured extraction to business software
- Moderate text and images with hosted safety models
- Deploy latency-sensitive inference for production applications