The AI arena is free today

Open Superagent
API Provider101 active models22 organizations

DeepInfra: API pricing, speed & models

DeepInfra hosts 101 active AI models, with input pricing from $0.02 per 1M tokens, with median throughput of 24 characters/sec, and P95 time to first token of 17.76s, with 98.1% success rate over 7 days. Compare DeepInfra's API speed, pricing, and reliability against other inference providers.

Median throughput24 c/s
Median TTFT2.15s
P95 TTFT17.76s
Success rate (7d)98.1%
From$0.02 /M tok

Catalog

Alibaba Cloud / Qwen Team39Google37Anthropic14DeepSeek14ByteDance12Meta11Moonshot AI8Zhipu AI8
Type
Price
189 models
Model
Qwen3.8 Maxfp4cache pricing
Qwen3.8-27Bcache pricing
Qwen3.8-27Bcache pricing
Qwen3.7 Maxcache pricing
Qwen3.5-35B-A3Bfp8cache pricing
Qwen3.5-35B-A3Bfp8cache pricing
Qwen3.5-35B-A3Bfp8cache pricing
Qwen3 MaxExclusivecache pricing
Qwen3.8 Maxcache pricing
At a glance

DeepInfrapricing, performance & catalog

The citable facts about DeepInfra's 101 models — sourced from provider APIs and refreshed continuously.

Lowest price
Mistral NeMo Instruct at $0.019 per 1M input tokens
Highest median throughput
DeepSeek-V4-Pro-0813 at 98 chars/s
Lowest median TTFT
MiMo-V2.5 at 1.09s
Largest context
DeepSeek-V4.1-Flash at 1.0M tokens
Catalog
101 active models from 22 organizations

FAQ

Common questions about DeepInfra.

What is DeepInfra?

DeepInfra is an API provider that hosts large language models. Active models: 101; From (input): $0.02 / 1M tok; Median throughput: 24 c/s; Median TTFT: 2.15s; Success rate (7d): 98.1%.

How many models does DeepInfra offer?

DeepInfra currently serves 101 active models out of 122 historical offerings on LLM Stats.

What is DeepInfra's API pricing?

DeepInfra input pricing starts from $0.02 per 1M tokens, with the most expensive offering at $10 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is DeepInfra?

DeepInfra delivers a median throughput of 24 characters per second and P95 time to first token of 17.76s across its model catalog. See the Models tab for per-model throughput and TTFT breakdowns.

Is DeepInfra reliable?

DeepInfra has a 98.1% success rate across 210 API calls in the last 7 days, with a 1.9% error rate.

Does DeepInfra support function calling?

Yes. 175 of 101 models on DeepInfra support function calling (tool use). The Capabilities tab lists which specific models accept tool definitions.

Does DeepInfra support JSON mode and structured output?

Yes. 164 of 101 models on DeepInfra support structured output (JSON mode / schema-constrained generation). The Capabilities tab shows which specific models accept response_format or json_schema parameters.

Does DeepInfra offer batch inference?

Yes. 186 of 101 models on DeepInfra support batch inference for cheaper, asynchronous workloads.

Does DeepInfra support multimodal models?

Yes. DeepInfra's catalog includes 52 vision-capable models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does DeepInfra host?

DeepInfra hosts models from Anthropic, ByteDance, DeepSeek, Google, Gryphe, and IBM, plus 16 more. See the Models tab for the full catalog grouped by creator.

How do I start using DeepInfra?

Sign up at https://deepinfra.com/ to get an API key, then call DeepInfra's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at DeepInfra's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.