The AI arena is free today

Open Playground
API Provider10 active models7 organizations

Together: API pricing, speed & models

Together hosts 10 active AI models, with input pricing from $0.30 per 1M tokens, with median throughput of 123 characters/sec, and P95 time to first token of 23.35s, with 53.5% success rate over 7 days. Compare Together's API speed, pricing, and reliability against other inference providers.

Median throughput123 c/s
Median TTFT1.57s
P95 TTFT23.35s
Success rate (7d)53.5%
From$0.30 /M tok

Catalog

Moonshot AI8Alibaba Cloud / Qwen Team6MiniMax3Google2DeepSeek1Zhipu AI1
Type
Price
21 models
At a glance

Togetherpricing, performance & catalog

The citable facts about Together's 10 models — sourced from provider APIs and refreshed continuously.

Lowest price
MiniMax M3 at $0.300 per 1M input tokens
Highest median throughput
Kimi K2.7 Code at 961 chars/s
Lowest median TTFT
Kimi K2.7 Code at 0.33s
Largest context
Kimi K3 at 1.0M tokens
Catalog
10 active models from 7 organizations

FAQ

Common questions about Together.

What is Together?

Together is an API provider that hosts large language models. Active models: 10; From (input): $0.30 / 1M tok; Median throughput: 123 c/s; Median TTFT: 1.57s; Success rate (7d): 53.5%.

How many models does Together offer?

Together currently serves 10 active models out of 24 historical offerings on LLM Stats.

What is Together's API pricing?

Together input pricing starts from $0.30 per 1M tokens, with the most expensive offering at $3 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is Together?

Together delivers a median throughput of 123 characters per second and P95 time to first token of 23.35s across its model catalog. See the Models tab for per-model throughput and TTFT breakdowns.

Is Together reliable?

Together has a 53.5% success rate across 157 API calls in the last 7 days, with a 46.5% error rate.

Does Together support function calling?

Yes. 21 of 10 models on Together support function calling (tool use). The Capabilities tab lists which specific models accept tool definitions.

Does Together support JSON mode and structured output?

Yes. 21 of 10 models on Together support structured output (JSON mode / schema-constrained generation). The Capabilities tab shows which specific models accept response_format or json_schema parameters.

Does Together offer batch inference?

Yes. 18 of 10 models on Together support batch inference for cheaper, asynchronous workloads.

Does Together support multimodal models?

Yes. Together's catalog includes 7 vision-capable models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does Together host?

Together hosts models from DeepSeek, Google, MiniMax, Moonshot AI, Alibaba Cloud / Qwen Team, and Zhipu AI, plus 1 more. See the Models tab for the full catalog grouped by creator.

How do I start using Together?

Sign up at https://together.ai/ to get an API key, then call Together's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at Together's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.