- Organizations
- thinking-machines
- Inkling-Small
Inkling-Small: API Pricing, Context Window & Benchmarks
Inkling-Small is a language model from Unknown Organization, released in July 2026, with multimodal input, a 256K-token context window, and pricing from $0.300/M input, $0.060/M cached input, $1.20/M output.
Inkling-Small is Thinking Machines Lab's efficient open-weights MoE multimodal model (276B total / 12B active parameters) released under Apache 2.0. It accepts text, image, and audio inputs and generates text, with native reasoning,
Inkling-Small benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
How Inkling-Small performs across real-world prompt categories.
Performance by conversation depth
How Inkling-Small holds up as conversations get longer.
Quality Tracker
Inkling-Small Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Inkling-Small pricing
Providers
Inkling-Small starts at $0.300 per million input tokens and $1.20 per million output tokens via Thinking Machines Lab. Reused prompt prefixes cost $0.0600 per million cached input tokens.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.300 | $0.0600 | $1.20 | 256.0K/256.0K | — | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Inkling-Small model size
Inkling-Small has 276 billion parameters. See how it compares to other models in the same parameter range.
Inkling-Small context window
Input and output token limits for Inkling-Small, plus how it ranks on long-context understanding.
Inkling-Small API
Available from the model provider
Inkling-Small is available from Thinking Machines Lab. It is not currently routed through the LLM Stats gateway.
Read the official API documentationInkling-Small latency
Inkling-Small time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Inkling-Small examples
Recent arena outputs from Inkling-Small, picked from the highest-ranked matchups.
Inkling-Small license
Inkling-Small is released under the Apache 2.0 license, which permits commercial use, has 276.0B parameters.
- License
- Apache 2.0
- Commercial use allowed
- Parameters
- 276.0B
Apache License 2.0 - allows commercial use
Inkling-Small resources
Official sources for Inkling-Small: api documentation, official playground, source repository, model weights.
Inkling-Small vs other models
The most-compared alternatives to Inkling-Small are MiMo-V2.5-Pro, GPT-5.2, Kimi K2.6. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Inkling-Small
Models ranked just above and below Inkling-Small by LLM Stats score.
FAQ
Common questions about Inkling-Small.