The AI arena is free today

Open Superagent
API Provider7 active models4 organizations

FriendliAI: API pricing, speed & models

FriendliAI hosts 7 active AI models, with input pricing from $0.14 per 1M tokens, with median throughput of 189 characters/sec, and P95 time to first token of 11.29s, with 98.2% success rate over 7 days. Compare FriendliAI's API speed, pricing, and reliability against other inference providers.

Median throughput189 c/s
Median TTFT1.26s
P95 TTFT11.29s
Success rate (7d)98.2%
From$0.14 /M tok

Catalog

Zhipu AI7Alibaba Cloud / Qwen Team3Google2
Type
Price
12 models
Model
GLM-5.3-Flashcache pricing
GLM-5.3-Flashcache pricing
GLM-5.3-Flashcache pricing
GLM-5.3cache pricing
GLM-5.2cache pricing
GLM-5.1cache pricing
At a glance

FriendliAIpricing, performance & catalog

The citable facts about FriendliAI's 7 models — sourced from provider APIs and refreshed continuously.

Lowest price
Gemma 4 31B at $0.140 per 1M input tokens
Highest median throughput
GLM-5.3-Flash at 204 chars/s
Lowest median TTFT
GLM-5.3-Flash at 1.21s
Largest context
GLM-5.3-Flash at 1.0M tokens
Catalog
7 active models from 4 organizations

FAQ

Common questions about FriendliAI.

What is FriendliAI?

FriendliAI is an API provider that hosts large language models. Active models: 7; From (input): $0.14 / 1M tok; Median throughput: 189 c/s; Median TTFT: 1.26s; Success rate (7d): 98.2%.

How many models does FriendliAI offer?

FriendliAI currently serves 7 active models out of 8 historical offerings on LLM Stats.

What is FriendliAI's API pricing?

FriendliAI input pricing starts from $0.14 per 1M tokens, with the most expensive offering at $1.4 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is FriendliAI?

FriendliAI delivers a median throughput of 189 characters per second and P95 time to first token of 11.29s across its model catalog. See the Models tab for per-model throughput and TTFT breakdowns.

Is FriendliAI reliable?

FriendliAI has a 98.2% success rate across 55 API calls in the last 7 days, with a 1.8% error rate.

Does FriendliAI support function calling?

Yes. 12 of 7 models on FriendliAI support function calling (tool use). The Capabilities tab lists which specific models accept tool definitions.

Does FriendliAI support JSON mode and structured output?

Yes. 12 of 7 models on FriendliAI support structured output (JSON mode / schema-constrained generation). The Capabilities tab shows which specific models accept response_format or json_schema parameters.

Does FriendliAI support multimodal models?

Yes. FriendliAI's catalog includes 3 vision-capable models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does FriendliAI host?

FriendliAI hosts models from Google, Alibaba Cloud / Qwen Team, Zhipu AI, and LG AI Research. See the Models tab for the full catalog grouped by creator.

How do I start using FriendliAI?

Sign up at https://friendli.ai/ to get an API key, then call FriendliAI's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at FriendliAI's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.