The AI arena is free today

Open Playground
API Provider12 active models8 organizations

Fireworks: API pricing, speed & models

Fireworks hosts 12 active AI models, with input pricing from $0.14 per 1M tokens, with median throughput of 93 characters/sec, and P95 time to first token of 8.39s, with 93.2% success rate over 7 days. Compare Fireworks's API speed, pricing, and reliability against other inference providers.

Fine-tuning
Median throughput93 c/s
Median TTFT1.65s
P95 TTFT8.39s
Success rate (7d)93.2%
From$0.14 /M tok

Catalog

Moonshot AI8MiniMax5Alibaba Cloud / Qwen Team3DeepSeek2Fireworks AI2Zhipu AI1
Type
Price
21 models
Model
Kimi K3cache pricing
Kimi K3cache pricing
At a glance

Fireworkspricing, performance & catalog

The citable facts about Fireworks's 12 models — sourced from provider APIs and refreshed continuously.

Lowest price
DeepSeek-V4-Flash-0731 at $0.140 per 1M input tokens
Highest median throughput
Qwen3.7-Plus at 981 chars/s
Lowest median TTFT
Kimi K2.6 at 0.34s
Largest context
DeepSeek-V4-Flash-0731 at 1.0M tokens
Catalog
12 active models from 8 organizations

FAQ

Common questions about Fireworks.

What is Fireworks?

Fireworks is an API provider that hosts large language models. Active models: 12; From (input): $0.14 / 1M tok; Median throughput: 93 c/s; Median TTFT: 1.65s; Success rate (7d): 93.2%.

How many models does Fireworks offer?

Fireworks currently serves 12 active models out of 39 historical offerings on LLM Stats.

What is Fireworks's API pricing?

Fireworks input pricing starts from $0.14 per 1M tokens, with the most expensive offering at $3 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is Fireworks?

Fireworks delivers a median throughput of 93 characters per second and P95 time to first token of 8.39s across its model catalog. See the Models tab for per-model throughput and TTFT breakdowns.

Is Fireworks reliable?

Fireworks has a 93.2% success rate across 836 API calls in the last 7 days, with a 6.8% error rate.

Does Fireworks support function calling?

Yes. 19 of 12 models on Fireworks support function calling (tool use). The Capabilities tab lists which specific models accept tool definitions.

Does Fireworks support JSON mode and structured output?

Yes. 17 of 12 models on Fireworks support structured output (JSON mode / schema-constrained generation). The Capabilities tab shows which specific models accept response_format or json_schema parameters.

Does Fireworks offer batch inference?

Yes. 19 of 12 models on Fireworks support batch inference for cheaper, asynchronous workloads.

Does Fireworks support fine-tuning?

Yes. 6 of 12 models on Fireworks support fine-tuning. The Capabilities tab lists which specific models accept fine-tuning jobs.

Does Fireworks support multimodal models?

Yes. Fireworks's catalog includes 5 vision-capable models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does Fireworks host?

Fireworks hosts models from DeepSeek, Fireworks AI, MiniMax, Moonshot AI, Alibaba Cloud / Qwen Team, and Zhipu AI, plus 2 more. See the Models tab for the full catalog grouped by creator.

How do I start using Fireworks?

Sign up at https://fireworks.ai/ to get an API key, then call Fireworks's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at Fireworks's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.