The AI arena is free today

Open Superagent

Sarvam-105B vs Step-3.5-Flash

Step-3.5-Flash leads the LLM Stats Score 37.2 to 25.7.

Sarvam AI · StepFun · Updated for 2026

Which is better?

Step-3.5-Flash leads the overall LLM Stats Score 37.2 to 25.7, ranking #72 overall.

In the 4 individual benchmarks reported for both models, Step-3.5-Flash wins 4; this is a narrower head-to-head signal than the composite indexes.

Based on current LLM Stats indexes, shared benchmarks, pricing, and model metadata for 2026.

Choose Sarvam-105B

  • you want the most recent training data — it shipped Mar 2026

Choose Step-3.5-Flash

  • overall performance matters — it scores 37.2 and ranks #72 on LLM Stats
  • your work emphasizes reasoning and coding — it leads those capability indexes
  • you value its reported benchmark strengths — it wins 4 of 4 exact shared results

At a glance

The differences that matter most.

Core performance indexes
25.7
#157
37.2
#72
26.2
#151
37.2
#72
-1.6
#251
20.5
#99
5.5
#148
15.3
#88
Cost, coverage & limits
Benchmark wins
0 of 4
4 of 4
Input price
— / M
$0.10 / M
Output price
— / M
$0.40 / M
Context window
65,536

Capability indexes

Additional strengths measured across groups of related public benchmarks

1 shared
Index
Sarvam-105B
Step-3.5-Flash
29.1#88
33.9#50
Conservative TrueSkill rating · higher is betterHow scores work

Individual benchmarks

14 reported for Sarvam-105B · 7 for Step-3.5-Flash

4 shared

Sarvam-105B outperforms in 0 benchmarks, while Step-3.5-Flash is better at 4 benchmarks (AIME 2025, BrowseComp, LiveCodeBench v6, SWE-Bench Verified).

Step-3.5-Flash significantly outperforms across most benchmarks.

Tue Sep 22 2026 • llm-stats.com

Human preference

Blind head-to-head votes and playground preference scores

Model Size

Parameter count comparison

91.0B diff

Step-3.5-Flash has 91.0B more parameters than Sarvam-105B, making it 86.7% larger.

Sarvam AI
Sarvam-105B
105.0Bparameters
StepFun
Step-3.5-Flash
196.0Bparameters
105.0B
Sarvam-105B
196.0B
Step-3.5-Flash

Context Window

Maximum input and output token capacity

Only Step-3.5-Flash specifies input context (65,536 tokens). Only Step-3.5-Flash specifies output context (8,192 tokens).

Sarvam AI
Sarvam-105B
Input- tokens
Output- tokens
StepFun
Step-3.5-Flash
Input65,536 tokens
Output8,192 tokens
Tue Sep 22 2026 • llm-stats.com

License

Usage and distribution terms

Both models are licensed under Apache 2.0.

Both models share the same licensing terms, providing consistent usage rights.

Sarvam-105B

Apache 2.0

Open weights

Step-3.5-Flash

Apache 2.0

Open weights

Release Timeline

When each model was launched

Sarvam-105B was released on 2026-03-06, while Step-3.5-Flash was released on 2026-02-02.

Sarvam-105B is 1 month newer than Step-3.5-Flash.

Sarvam-105B

Mar 6, 2026

6 months ago

1mo newer
Step-3.5-Flash

Feb 2, 2026

7 months ago

Knowledge Cutoff

When training data ends

Neither model specifies a knowledge cutoff date.

Unable to compare the recency of their training data.

No cutoff dates available

Outputs Comparison

Notice missing or incorrect data?Start an Issue discussion

Judge for yourself.

Run your own prompts against Sarvam-105B and Step-3.5-Flash side-by-side, then vote on the output you prefer.

Sarvam-105B
✓ Preferred
Step-3.5-Flash
Open in Playground

FAQ

Common questions about Sarvam-105B vs Step-3.5-Flash.

Which is better, Sarvam-105B or Step-3.5-Flash?

Step-3.5-Flash leads the LLM Stats Score 37.2 to 25.7. Sarvam-105B is made by Sarvam AI and Step-3.5-Flash is made by StepFun. The best choice depends on your use case — compare their capability indexes, individual benchmarks, pricing, and limits above.

How does Sarvam-105B compare to Step-3.5-Flash in benchmarks?

Sarvam-105B scores MATH-500: 98.6%, AIME 2025: 96.7%, MMLU: 90.6%, HMMT 2025: 85.8%, HMMT25: 85.8%. Step-3.5-Flash scores AIME 2025: 97.3%, Tau-bench: 88.2%, LiveCodeBench v6: 86.4%, IMO-AnswerBench: 85.4%, SWE-Bench Verified: 74.4%.

What are the context window sizes for Sarvam-105B and Step-3.5-Flash?

Sarvam-105B supports an unknown number of tokens and Step-3.5-Flash supports 66K tokens. A larger context window lets you process longer documents, conversations, or codebases in a single request.

What are the main differences between Sarvam-105B and Step-3.5-Flash?

Key differences include LLM Stats Score (25.7 vs 37.2). See the full comparison above for benchmark-by-benchmark results.

Who makes Sarvam-105B and Step-3.5-Flash?

Sarvam-105B is developed by Sarvam AI and Step-3.5-Flash is developed by StepFun.