HumeAI/voice-controllability-leaderboard
Voice Controllability Leaderboard
Capability and preference boards for voice design and instruction following, rated by paid blind human raters. Hume systems are hidden from every board; unreleased systems are hidden from the Role Fit boards (their matchups still count toward shown models' win rates).
Scores come from three sources: the humeval voice_creation evals (how well raters can hear the requested attribute, plus directly measured prosody dials), the inline style-tags eval, and the VCP human study (blind A/B votes).
Tabs:
- Voice Design: voice_creation vc-mode evals — an Overall pill (the five category scores side by side with their mean) over one section per attribute (accent, age, gender, texture, and the directly measured prosody dials), vc-capable models
- Multilingual: the same vc roster scored by native raters across 13 languages — 1-5 likerts rather than accuracy against an answer key, so it is kept out of the Voice Design mean. Overall plus one column per question, with by-language (a table of language codes) and by-accent breakouts
- Instruct: voice_creation instruct-mode evals (Tone: delivery overall + 5 subcategories; Emotion section: overall + per-emotion identification), instruct-capable models
- Style Tags: the inline style-tags eval (Single Tag, Sequential, and a by-tag-type breakout)
- Role Fit: VCP human-study pairwise win rates (role fit and naturalness topline, plus a per-use-case win-rate heatmap matrix)
The Voice Design, Instruct, and Role Fit tabs pair a grouped-bar SVG chart (baked into the board JSON by build_boards.py, reusing analysis/controllability/charts.py) with the heatmap table.
Sibling of the Real World VoiceEQ Benchmark Space, sharing its app machinery and board-JSON schema, on its own private dataset.
Boards are built from the humeval repo, which holds the evals, the build scripts and the maintenance notes.
