At a glance

Groqpricing, performance & catalog

The citable facts about Groq's 4 models — sourced from provider APIs and refreshed continuously.

Lowest price
GPT OSS 120B at $0.150 per 1M input tokens
Highest throughput
GPT OSS 120B at 500 tokens/s
Lowest latency
GPT OSS 120B at 0.50s
Largest context
Whisper Large V3 Turbo at 100.0M tokens
Catalog
4 active models from 3 organizations

FAQ

Common questions about Groq.

What is Groq?

Groq is an API provider that hosts large language models. Active models: 4; From (input): $0.15 / 1M tok; Avg throughput: 500 tok/s; Avg latency: 0.50 s; Max context: 100.0M.

How many models does Groq offer?

Groq currently serves 4 active models out of 11 historical offerings on LLM Stats.

What is Groq's API pricing?

Groq input pricing starts from $0.15 per 1M tokens, with the most expensive offering at $0.15 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is Groq?

Groq averages 500 output tokens per second across its catalog, with average latency of 0.50s. Per-model performance is shown in the Performance tab.

Is Groq OpenAI compatible?

Most providers expose an OpenAI-compatible /v1/chat/completions endpoint so you can switch from OpenAI to Groq by changing only the base URL and API key. Check https://groq.com/ for the exact endpoint format and any provider-specific parameters.

Does Groq support multimodal models?

Yes. Groq's catalog includes 1 audio models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does Groq host?

Groq hosts models from OpenAI, PlayAI, and Meta. See the Models tab for the full catalog grouped by creator.

How do I start using Groq?

Sign up at https://groq.com/ to get an API key, then call Groq's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at Groq's base URL with your key. Use the Pricing and Performance tabs above to pick the right model for your latency, cost, and context-window requirements.