datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tadabur-align-references
tadabur-align-references
Precomputed reference embeddings powering tadabur-align — word-level timestamp extraction for Quranic recitation via DTW alignment transfer (no ASR).
What this is
For 5,481 of the Quran's 6,236 ayahs, this dataset holds frame-level tadabur-embedding features for up to 8 reference reciters, plus each reference's word-level timestamps and internal-pause intervals. No audio is included — only model outputs and timing data. tadabur-align… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur-align-references.moss-character-reference-voices
MOSS character reference voices (1336 voices)
1336 distinct synthetic character voices, each mined from a cluster of generated MOSS-VA-v2 character
audio and auto-annotated by Gemini-3-Flash. For every cluster the model was shown the 3 cluster samples
their automatic voice scores, chose the single most representative sample, and wrote a full
casting-style profile.
Contents
dataset.jsonl — one row per voice: cid, name, tagline, description, age, gender, register… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-character-reference-voices.asr-reference-set-eval-temp
Temporary ASR evaluation audio
Temporary public audio files used for hosted ASR evaluation.
moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.Celtic_Stems_Reference_Sessions_Preview
Harmonic Frontier Audio – Celtic Constellation Reference Sessions (Preview, v0.9)
A high-fidelity music-production dataset designed to connect isolated source performances, production processing, arrangement context, and finished musical outcomes.
Celtic Constellation Reference Sessions (Preview), created by Harmonic Frontier Audio, introduces the Reference Sessions product vertical through a compact proof-of-concept built around purpose-recorded Celtic ensemble material.… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Celtic_Stems_Reference_Sessions_Preview.reference-logitsThis is just a temp scratchpad for sharing data.
hindi-emotion-voice-references
Emotion Voice References for QwenTTS
Emotion-wise organized voice WAV files for Hindi (Male + Female) and English (Male + Female), ready to use as reference audio for Qwen3-TTS voice cloning.
🎯 Purpose
Reference audio clips for Qwen3-TTS voice cloning with emotion-aware synthesis. Organized consistently as {language}/{gender}/{emotion}/*.wav.
📁 Structure
hindi/
male/{emotion}/ → 4,895 WAVs (HIN_M_* from Rasa Hindi + URD_M_* from Rasa Urdu)… See the full description on the dataset page: https://huggingface.co/datasets/sumedhu/hindi-emotion-voice-references.6k-diverse-reference-voices
6k Diverse Reference Voices
6,064 permissively licensed reference voices for casting expressive voice-acting generations.
All voices in this collection are permissively usable: they were either synthetically created or
extracted from the CC-BY part of Emilia. Licensed under CC-BY-4.0.
Source / attribution: derived from TTS-AGI/moss-reference-voices-consolidated (CC-BY-4.0),
re-published under LAION with clarified metadata documentation. If you use this dataset,
please attribute… See the full description on the dataset page: https://huggingface.co/datasets/laion/6k-diverse-reference-voices.Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.reference-voices-enhanced
Reference Voices Enhanced
2,004 AI voice samples enhanced with ClearerVoice-Studio MossFormer2_SE_48K speech enhancement, annotated with Empathic Insight Voice Plus (59 quality + emotion scores).
Dataset Summary
Source: laion/ai-voices-deduplicated (2,004 speaker-deduplicated, quality-filtered AI voice samples)
Speech Enhancement: ClearerVoice MossFormer2_SE_48K — background noise removal and speech clarity improvement
Output Format: Enhanced WAV files at 48kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/reference-voices-enhanced.voice_referencesmoss-reference-voices-consolidated
MOSS reference voices — consolidated (6,064 voices)
6,064 reference voices for casting MOSS-VA-v2 voice-acting generations. Every voice was auto-annotated
by Gemini (name, tagline, language, accent, age/gender read, register, timbre, distinctive features,
emotional range, casting suggestions for 4 genres, free-text tags, search text) and scored on 99 measured
dimensions: 57 VoiceNet voice-quality axes (timbre/prosody/register/speaking-style, e.g. brightness,
roughness, warmth… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-reference-voices-consolidated.voice-referenceclap-reference-librariesen_and_de_reference_voice_files_for_emotion_cloningml2021-reference-corpusEmotion-Voice-Attribute-Reference-Snippets-DACVAE
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: ~98 GB (latents-only, no WAV)
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.nptel_indian_en_subset_with_reference_v1german_premium_tts_referencenptel_indian_en_subset_with_reference_v0abico-referencebeethoven-reference-audioemotion_dataset_for_tts_with_transcriptions_and_reference_voice_v1guzman-clone-voice-reference
