datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
talent_plus_rl_groups_of_50_with_audiobox_scoresrelease-scores
TuneJury Reward Scores
Pre-computed TuneJury reward scores for seven open-license music collections (219,020 clips total). Companion artifact to the paper TuneJury: An Open Metric for Improving Music Generation Preference Alignment (arXiv:2606.17006, code).
This dataset ships scores and identifiers only, not audio. Each row is one deterministic TuneJury scorer call per clip. To obtain the audio, fetch each collection from its original source (see "Audio sources" below).
The… See the full description on the dataset page: https://huggingface.co/datasets/TuneJury/release-scores.nollywood-ctc-scored-ep3-hauwavocal-score-synthetic-smoke-test
Vocal Score Synthetic Smoke Test
A tiny, fully synthetic fixture for smoke-testing monophonic vocal-to-notation pipelines. It contains one ten-second “ah”-like synthesized rendition of the public-domain melody commonly known as “Twinkle, Twinkle, Little Star,” its scripted note sequence, and observed exports from my Vocal Score research pipeline.
No human voice, copyrighted recording, voice embedding, lyric, or personal data is included.
Files
twinkle.wav — 22.05… See the full description on the dataset page: https://huggingface.co/datasets/mattwinwood/vocal-score-synthetic-smoke-test.banger-scorer-generated-songs
Banger Scorer Generated Songs
230 AI-generated songs across 10 genres and 5 languages, all scored by the banger scorer. Includes 10 genre tests (20 songs each) plus 1 banger-optimized run (30 songs) that used data-driven parameter selection to maximize scores.
Every song includes its MP3 audio, generation metadata (BPM, key, seed, caption prompt), and banger score. Useful for research into AI music quality, training better scorers, or just listening to what works and what does not.… See the full description on the dataset page: https://huggingface.co/datasets/treadon/banger-scorer-generated-songs.transcription-scorer
Transcription Scorer Dataset
The Transcription Scorer dataset was created to support research in reference-free evaluation of Automatic Speech Recognition (ASR) systems using human feedback. Unlike traditional evaluation metrics such as WER and its derivatives, this dataset reflects judgments of ASR outputs by human raters across multiple criteria, simulating the way a teacher grades students.
⚙️ What’s Inside
This dataset contains 1200 audio samples (from diverse sources… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/transcription-scorer.spoken-web-questions-scoretrivia_qa-audio-scorellama-questions-scorespoken-web-questions-text-scoredeepspeech_with_qwen_description_exp1_score_with_emotion_and_weryoutube-kinyarwanda-snac-scoredwur_ctc_kln_scoredspoken-web-questions-text_original-scoreaudio_L2-regular_llama-questions-scoretrivia_qa-audio-text_original-scorecheckpoint_similarity_scores_with_audiotrivia_qa-audio-text-scoretrivia_qa-audio-ASR_GT-scoreiemocap-emotion-scoresllama-questions-text-scorellama-questions-text_original-scorespoken-web-questions-ASR_GT-scoredeepspeech_with_qwen_description_exp1_scorellama-questions-ASR_GT-scoreaudio_L2-regular_spoken-web-questions-scoreaudio_L2-regular_trivia_qa-audio-score
