The AI arena is free today

Open Superagent
Unknown OrganizationReleased on Jul 30, 2026

Inkling-Small: API Pricing, Context Window & Benchmarks

Inkling-Small is a language model from Unknown Organization, released in July 2026, with multimodal input, a 256K-token context window, and pricing from $0.300/M input, $0.060/M cached input, $1.20/M output.

Inkling-Small is Thinking Machines Lab's efficient open-weights MoE multimodal model (276B total / 12B active parameters) released under Apache 2.0. It accepts text, image, and audio inputs and generates text, with native reasoning,

Input
TextImageAudio
Output
Text

Inkling-Small benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

How Inkling-Small performs across real-world prompt categories.

Performance by conversation depth

How Inkling-Small holds up as conversations get longer.

Quality Tracker

Inkling-Small Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Tue Aug 25 2026
Notice missing or incorrect data?

Inkling-Small pricing

Providers

Inkling-Small starts at $0.300 per million input tokens and $1.20 per million output tokens via Thinking Machines Lab. Reused prompt prefixes cost $0.0600 per million cached input tokens.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
Thinking Machines Lab logoThinking Machines Lab
$0.300$0.0600$1.20256.0K/256.0K
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Inkling-Small model size

Inkling-Small has 276 billion parameters. See how it compares to other models in the same parameter range.

Parameters
276BMoE
Frontier (200B+)
276B
1B7B70B405B

Inkling-Small context window

Input and output token limits for Inkling-Small, plus how it ranks on long-context understanding.

InputOutput
256Ktokens
256Ktokens
385 pages of text
256K
8K128K1M

Inkling-Small API

Available from the model provider

Inkling-Small is available from Thinking Machines Lab. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Inkling-Small latency

Inkling-Small time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Inkling-Small examples

Recent arena outputs from Inkling-Small, picked from the highest-ranked matchups.

Inkling-Small license

Inkling-Small is released under the Apache 2.0 license, which permits commercial use, has 276.0B parameters.

License
Apache 2.0
Commercial use allowed
Parameters
276.0B

Apache License 2.0 - allows commercial use

Inkling-Small resources

Official sources for Inkling-Small: api documentation, official playground, source repository, model weights.

Inkling-Small vs other models

The most-compared alternatives to Inkling-Small are MiMo-V2.5-Pro, GPT-5.2, Kimi K2.6. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Inkling-Small

Models ranked just above and below Inkling-Small by LLM Stats score.

 

MiMo-V2.5-Pro

Score pending
 

GPT-5.2

Score pending
 

Kimi K2.6

Score pending
 

Gemma 4 26B-A4B

Score pending
 

Solar Pro 4

Score pending
 

Qwen3.8-27B

Score pending

FAQ

Common questions about Inkling-Small.

When was Inkling-Small released?

Inkling-Small was released on July 30, 2026 by thinking-machines. This is the official Inkling-Small release date tracked on LLM Stats.

How much does Inkling-Small cost?

Inkling-Small pricing starts at $0.30 per million input tokens, $0.06 per million cached input tokens, $1.20 per million output tokens via Thinking Machines Lab, the lowest price among tracked providers.

Is Inkling-Small available via API?

Yes, Inkling-Small is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

How big is Inkling-Small?

Inkling-Small has 276 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Inkling-Small?

Inkling-Small was created by thinking-machines.

What is the license for Inkling-Small?

Inkling-Small is released under the Apache 2.0 license. This is an open-source / open-weight license that permits self-hosting.

Is Inkling-Small multimodal?

Yes, Inkling-Small is multimodal and can accept both text and images as input.

Where can I use Inkling-Small?

Inkling-Small is available through 1 provider including Thinking Machines Lab.

Where is the Inkling-Small paper or technical report?

Inkling-Small has a paper or technical report available at https://thinkingmachines.ai/news/inkling-small/. Use that source for architecture, training, release and evaluation details.

What models should I compare Inkling-Small against?

Common Inkling-Small comparisons include Inkling-Small vs MiMo-V2.5-Pro, Inkling-Small vs GPT-5.2, Inkling-Small vs Kimi K2.6. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.