CoolFace
Apppublic

HumeAI/voice-controllability-leaderboard

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
1likes
App README

Voice Controllability Leaderboard

Capability and preference boards for voice design and instruction following, rated by paid blind human raters. Hume systems are hidden from every board; unreleased systems are hidden from the Role Fit boards (their matchups still count toward shown models' win rates).

Scores come from three sources: the humeval voice_creation evals (how well raters can hear the requested attribute, plus directly measured prosody dials), the inline style-tags eval, and the VCP human study (blind A/B votes).

Tabs:

  • —Voice Design: voice_creation vc-mode evals — an Overall pill (the five category scores side by side with their mean) over one section per attribute (accent, age, gender, texture, and the directly measured prosody dials), vc-capable models
  • —Multilingual: the same vc roster scored by native raters across 13 languages — 1-5 likerts rather than accuracy against an answer key, so it is kept out of the Voice Design mean. Overall plus one column per question, with by-language (a table of language codes) and by-accent breakouts
  • —Instruct: voice_creation instruct-mode evals (Tone: delivery overall + 5 subcategories; Emotion section: overall + per-emotion identification), instruct-capable models
  • —Style Tags: the inline style-tags eval (Single Tag, Sequential, and a by-tag-type breakout)
  • —Role Fit: VCP human-study pairwise win rates (role fit and naturalness topline, plus a per-use-case win-rate heatmap matrix)

The Voice Design, Instruct, and Role Fit tabs pair a grouped-bar SVG chart (baked into the board JSON by build_boards.py, reusing analysis/controllability/charts.py) with the heatmap table.

Sibling of the Real World VoiceEQ Benchmark Space, sharing its app machinery and board-JSON schema, on its own private dataset.

Boards are built from the humeval repo, which holds the evals, the build scripts and the maintenance notes.