- Providers
- DeepInfra
DeepInfra: API pricing, speed & models
DeepInfra hosts 101 active AI models, with input pricing from $0.02 per 1M tokens, with median throughput of 24 characters/sec, and P95 time to first token of 17.76s, with 98.1% success rate over 7 days. Compare DeepInfra's API speed, pricing, and reliability against other inference providers.
Median throughput24 c/s
Median TTFT2.15s
P95 TTFT17.76s
Success rate (7d)98.1%
From$0.02 /M tok
Catalog
Alibaba Cloud / Qwen Team39Google37Anthropic14DeepSeek14ByteDance12Meta11Moonshot AI8Zhipu AI8
Type
Price
189 models| Model |
|---|
Qwen3.8 Maxfp4cache pricing |
Qwen3.8-27Bcache pricing |
Qwen3.8-27Bcache pricing |
Qwen3.7 Maxcache pricing |
Qwen3.5-397B-A17Bfp8cache pricing |
Qwen3.5-397B-A17Bfp8cache pricing |
Qwen3.5-397B-A17Bfp8cache pricing |
Qwen3.5-35B-A3Bfp8cache pricing |
Qwen3.5-35B-A3Bfp8cache pricing |
Qwen3.5-35B-A3Bfp8cache pricing |
Qwen3 Max ThinkingExclusivecache pricing |
Qwen3 MaxExclusivecache pricing |
Qwen3-Coder 480B A35B Instructfp4Exclusivecache pricing |
Qwen3.8 Maxcache pricing |
Qwen3 VL 235B A22B Instructfp8cache pricing |
Qwen3 VL 235B A22B Instructfp8cache pricing |