The AI arena is free today

Open Superagent

VoiceBench Avg

Paper

Progress Over Time

Interactive timeline showing model performance evolution on VoiceBench Avg

State-of-the-art frontier
Open
Proprietary

VoiceBench Avg Leaderboard

2 models
ContextCostLicense
1
Thinking Machines Lab
Thinking Machines Lab
276B256K$0.30 / $1.20
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
7B
Notice missing or incorrect data?
About this benchmark

What is VoiceBench Avg?

VoiceBench is the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants, evaluating capabilities including general knowledge, instruction-following, reasoning, and safety using both synthetic and real spoken instruction data with diverse speaker characteristics and environmental conditions.

VoiceBench Avg is a multimodal benchmark evaluating models on chat, reasoning, safety, speech to text, general, and communication tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.

Compare leaders on the best AI for chat, best AI for reasoning, best AI for safety, best AI for speech to text, best AI for general and best AI for communication leaderboards.

Current leaders

Inkling-Small from Thinking Machines Lab currently leads the VoiceBench Avg leaderboard with a score of 0.901 across 2 evaluated AI models.

1Inkling-SmallThinking Machines Lab90.1%
2Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team74.1%

Source paper

Title
VoiceBench: Benchmarking LLM-Based Voice Assistants
Authors
Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao, and 2 others
Published
Abstract

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience compared to traditional text-based interactions. However, the absence of benchmarks designed to evaluate these speech interaction capabilities has hindered progress of LLM-based voice assistants development. Current evaluations focus primarily on automatic speech recognition (ASR) or general knowledge evaluation with clean speeches, neglecting the more intricate, real-world scenarios that involve diverse speaker characteristics, environmental and content factors. To address this, we introduce VoiceBench, the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants. VoiceBench also includes both real and synthetic spoken instructions that incorporate the above three key real-world variations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.

FAQ

Common questions about the VoiceBench Avg benchmark and leaderboard.

What is the VoiceBench Avg benchmark?

VoiceBench is the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants, evaluating capabilities including general knowledge, instruction-following, reasoning, and safety using both synthetic and real spoken instruction data with diverse speaker characteristics and environmental conditions.

What is the VoiceBench Avg leaderboard?

The VoiceBench Avg leaderboard ranks 2 AI models based on their performance on this benchmark. Currently, Inkling-Small by Thinking Machines Lab leads with a score of 0.901. The average score across all models is 0.821.

What is the highest VoiceBench Avg score?

The highest VoiceBench Avg score is 0.901, achieved by Inkling-Small from Thinking Machines Lab.

How many models are evaluated on VoiceBench Avg?

2 models have been evaluated on the VoiceBench Avg benchmark, with 0 verified results and 2 self-reported results.

Where can I find the VoiceBench Avg paper?

The VoiceBench Avg paper is available at https://arxiv.org/abs/2410.17196. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does VoiceBench Avg cover?

VoiceBench Avg is categorized under chat, reasoning, safety, speech to text, general, and communication. The benchmark evaluates multimodal models.

What is the best open-source model on VoiceBench Avg?

Inkling-Small by Thinking Machines Lab is the top-ranked open-source model on VoiceBench Avg, with a score of 0.901 (rank #1).

Which model offers the best value on VoiceBench Avg?

Among models scoring within 10% of the leader, Inkling-Small from Thinking Machines Lab is the cheapest, at $0.30 per million input tokens with a score of 0.901.

How recent are the VoiceBench Avg leaderboard results?

The VoiceBench Avg leaderboard was last updated in August 2026 and currently includes 2 evaluated models.