At a glance

xAIpricing, performance & catalog

The citable facts about xAI's 10 models — sourced from provider APIs and refreshed continuously.

Lowest price
Grok-4 Fast Non-Reasoning at $0.200 per 1M input tokens
Highest throughput
Grok-4 Fast Reasoning at 95 chars/s
Lowest latency
Grok-4 Fast Reasoning at 0.87s
Largest context
Grok-4 Fast Non-Reasoning at 2.0M tokens
Catalog
10 active models from 1 organization

FAQ

Common questions about xAI.

What is xAI?

xAI is an API provider that hosts large language models. Active models: 10; From (input): $0.20 / 1M tok; Median throughput: 90 c/s; P95 latency: 1.34s; Success rate (7d): 96%.

How many models does xAI offer?

xAI currently serves 10 active models out of 20 historical offerings on LLM Stats.

What is xAI's API pricing?

xAI input pricing starts from $0.20 per 1M tokens, with the most expensive offering at $3 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is xAI?

xAI delivers a median throughput of 90 characters per second and P95 latency of 1.34s across its model catalog. See the Models tab for per-model throughput and latency breakdowns.

Is xAI reliable?

xAI has a 96% success rate across 224 API calls in the last 7 days, with a 4% error rate.

Is xAI OpenAI compatible?

Most providers expose an OpenAI-compatible /v1/chat/completions endpoint so you can switch from OpenAI to xAI by changing only the base URL and API key. Check https://docs.x.ai for the exact endpoint format and any provider-specific parameters.

Does xAI support multimodal models?

Yes. xAI's catalog includes 9 vision-capable, 2 image generation, and 6 video models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does xAI host?

xAI hosts models from xAI. See the Models tab for the full catalog grouped by creator.

How do I start using xAI?

Sign up at https://docs.x.ai to get an API key, then call xAI's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at xAI's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.