API Provider20 active models13 organizations

Replicate: API pricing, speed & models

Replicate hosts 20 active AI models, with 78.2% success rate over 7 days. Compare Replicate's API speed, pricing, and reliability against other inference providers.

Image generationAudioVideo
Success rate (7d)78.2%

Catalog

Kling AI24Black Forest Labs8Sourceful8WAN Video8ByteDance6Alibaba4MiniMax2Alibaba Cloud / Qwen Team1
Type
Price
63 models
Model
At a glance

Replicatepricing, performance & catalog

The citable facts about Replicate's 20 models — sourced from provider APIs and refreshed continuously.

Largest context
Riverflow 2.0 Pro at 10K tokens
Catalog
20 active models from 13 organizations

Most affordable

No public pricing data.

Fastest median throughput

No throughput data yet.

FAQ

Common questions about Replicate.

What is Replicate?

Replicate is an API provider that hosts large language models. Active models: 20; Success rate (7d): 78.2%.

How many models does Replicate offer?

Replicate currently serves 20 active models out of 39 historical offerings on LLM Stats.

Is Replicate reliable?

Replicate has a 78.2% success rate across 234 API calls in the last 7 days, with a 21.8% error rate.

Does Replicate support multimodal models?

Yes. Replicate's catalog includes 29 vision-capable, 27 image generation, 12 audio, and 24 video models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does Replicate host?

Replicate hosts models from Alibaba, Black Forest Labs, ByteDance, Kling AI, MiniMax, and Alibaba Cloud / Qwen Team, plus 7 more. See the Models tab for the full catalog grouped by creator.

How do I start using Replicate?

Sign up at https://replicate.com/ to get an API key, then call Replicate's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at Replicate's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.