The AI arena is free today

Open Playground
API Provider13 active models5 organizations

Google: API pricing, speed & models

Google hosts 13 active AI models, with input pricing from $0.25 per 1M tokens, with median throughput of 277 characters/sec, and P95 time to first token of 56.42s, with 99.6% success rate over 7 days. Compare Google's API speed, pricing, and reliability against other inference providers.

Image generationVideo
Median throughput277 c/s
Median TTFT8.45s
P95 TTFT56.42s
Success rate (7d)99.6%
From$0.25 /M tok

Catalog

Type
Price
56 models
Model
Gemini 3 Flashcache pricing
Gemini 3 Flashcache pricing
Gemini 3 Flashcache pricing
Gemini 3 Flashcache pricing
At a glance

Googlepricing, performance & catalog

The citable facts about Google's 13 models — sourced from provider APIs and refreshed continuously.

Lowest price
Gemini 3.1 Flash-Lite at $0.250 per 1M input tokens
Highest median throughput
Gemini 3.5 Flash-Lite at 618 chars/s
Lowest median TTFT
Gemini 3 Flash at 2.27s
Largest context
Gemini Omni Flash at 1.0M tokens
Catalog
13 active models from 5 organizations

FAQ

Common questions about Google.

What is Google?

Google is an API provider that hosts large language models. Active models: 13; From (input): $0.25 / 1M tok; Median throughput: 277 c/s; Median TTFT: 8.45s; Success rate (7d): 99.6%.

How many models does Google offer?

Google currently serves 13 active models out of 44 historical offerings on LLM Stats.

What is Google's API pricing?

Google input pricing starts from $0.25 per 1M tokens, with the most expensive offering at $2.5 per 1M tokens. See the Pricing tab above for the full per-model breakdown.

How fast is Google?

Google delivers a median throughput of 277 characters per second and P95 time to first token of 56.42s across its model catalog. See the Models tab for per-model throughput and TTFT breakdowns.

Is Google reliable?

Google has a 99.6% success rate across 9.0K API calls in the last 7 days, with a 0.4% error rate.

Does Google support function calling?

Yes. 28 of 13 models on Google support function calling (tool use). The Capabilities tab lists which specific models accept tool definitions.

Does Google support JSON mode and structured output?

Yes. 28 of 13 models on Google support structured output (JSON mode / schema-constrained generation). The Capabilities tab shows which specific models accept response_format or json_schema parameters.

Does Google offer batch inference?

Yes. 18 of 13 models on Google support batch inference for cheaper, asynchronous workloads.

Does Google support multimodal models?

Yes. Google's catalog includes 20 vision-capable, 10 image generation, and 10 video models. See the Models and Capabilities tabs for the full per-model breakdown.

Whose models does Google host?

Google hosts models from Google, AI21 Labs, Anthropic, Meta, and Mistral AI. See the Models tab for the full catalog grouped by creator.

How do I start using Google?

Sign up at https://ai.google.dev to get an API key, then call Google's API directly from your application. Most clients work out of the box by pointing the OpenAI SDK at Google's base URL with your key. Use the Models and Pricing tabs above to pick the right model for your latency, cost, and context-window requirements.