The AI arena is free today

Open Playground
API Provider20 active models10 organizations

Novita: API pricing, speed & models

Novita hosts 20 active AI models, with input pricing from $0.10 per 1M tokens, with median throughput of 135 characters/sec, and P95 time to first token of 5.96s, with 92.7% success rate over 7 days. Compare Novita's API speed, pricing, and reliability against other inference providers.

Median throughput135 c/s
Median TTFT2.44s
P95 TTFT5.96s
Success rate (7d)92.7%
From$0.10 /M tok

Catalog

Alibaba Cloud / Qwen Team11Moonshot AI9MiniMax5Google4Xiaomi4DeepSeek3Zhipu AI1
Type
Price
37 models
Model
Qwen3.6 Pluscache pricing
Qwen3.6 Pluscache pricing
Qwen3.6 Pluscache pricing
Qwen3 32B32.43 QPS
Qwen3 30B A3B88.84 QPS
At a glance

Novitapricing, performance & catalog

The citable facts about Novita's 20 models — sourced from provider APIs and refreshed continuously.

Lowest price
Qwen3 32B at $0.100 per 1M input tokens
Highest median throughput
MiniMax M3 at 285 chars/s
Lowest median TTFT
MiMo-V2.5 at 1.03s
Largest context
DeepSeek-V4-Flash-0731 at 1.0M tokens
Catalog
20 active models from 10 organizations

FAQ

Common questions about Novita.

What is Novita?

Novita is an API provider that hosts large language models. Active models: 20; From (input): $0.10 / 1M tok; Median throughput: 135 c/s; Median TTFT: 2.44s; Success rate (7d): 92.7%.

How many models does Novita offer?

Novita currently serves 20 active models out of 52 historical offerings on LLM Stats.

What is Novita's API pricing?

Novita input pricing starts from $0.10 per 1M tokens, with the most expensive offering at $3 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is Novita?

Novita delivers a median throughput of 135 characters per second and P95 time to first token of 5.96s across its model catalog. See the Models tab for per-model throughput and TTFT breakdowns.

Is Novita reliable?

Novita has a 92.7% success rate across 1.9K API calls in the last 7 days, with a 7.3% error rate.

Does Novita support function calling?

Yes. 37 of 20 models on Novita support function calling (tool use). The Capabilities tab lists which specific models accept tool definitions.

Does Novita support JSON mode and structured output?

Yes. 34 of 20 models on Novita support structured output (JSON mode / schema-constrained generation). The Capabilities tab shows which specific models accept response_format or json_schema parameters.

Does Novita offer batch inference?

Yes. 24 of 20 models on Novita support batch inference for cheaper, asynchronous workloads.

Does Novita support multimodal models?

Yes. Novita's catalog includes 10 vision-capable models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does Novita host?

Novita hosts models from DeepSeek, Google, MiniMax, Moonshot AI, Alibaba Cloud / Qwen Team, and Xiaomi, plus 4 more. See the Models tab for the full catalog grouped by creator.

How do I start using Novita?

Sign up at https://novita.ai/ to get an API key, then call Novita's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at Novita's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.