MusicCaps
Progress Over Time
Interactive timeline showing model performance evolution on MusicCaps
MusicCaps Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 7B | — | — |
What is MusicCaps?
MusicCaps is a dataset composed of 5,521 music examples, each labeled with an English aspect list and a free text caption written by musicians. The dataset contains 10-second music clips from AudioSet paired with rich textual descriptions that capture sonic qualities and musical elements like genre, mood, tempo, instrumentation, and rhythm. Created to support research in music-text understanding and generation tasks.
MusicCaps is a multimodal benchmark evaluating models on multimodal and audio tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.3.
Compare leaders on the best AI for multimodal and best AI for audio leaderboards.
Current leaders
Qwen2.5-Omni-7B from Alibaba Cloud / Qwen Team currently leads the MusicCaps leaderboard with a score of 0.328 across 1 evaluated AI models.
Source paper
- Title
- MusicLM: Generating Music From Text
- Authors
- Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, and 9 others
- Published
- arXiv
- 2301.11325
Abstract
We introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff". MusicLM casts the process of conditional music generation as a hierarchical sequence-to-sequence modeling task, and it generates music at 24 kHz that remains consistent over several minutes. Our experiments show that MusicLM outperforms previous systems both in audio quality and adherence to the text description. Moreover, we demonstrate that MusicLM can be conditioned on both text and a melody in that it can transform whistled and hummed melodies according to the style described in a text caption. To support future research, we publicly release MusicCaps, a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts.
FAQ
Common questions about the MusicCaps benchmark and leaderboard.