datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laions_got_talent_enhanced_no_metadataEnhancementDetection_LibrittsTrainClean360Wham
Dataset Card for "EnhancementDetection_LibrittsTrainClean360Wham"
More Information needed
voxforge_spanish_enhanced
VoxForge Spanish Enhanced (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of the Spanish subset of:
VoxForge.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been satisfactory.
It has been created to… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/voxforge_spanish_enhanced.th-en-zh-tts-200k-enhanced
TH-EN-ZH Multi-speaker TTS Dataset (200K, RE-USE Enhanced)
Speech-enhanced variant of FILM6912/th-en-zh-tts-200k.
Every clip has been processed through NVIDIA RE-USE (universal speech enhancement, SEMamba) at its native sample rate, then re-encoded losslessly as FLAC (PCM_16).
Same schema, same row order, same 200,000 rows (th 100k / en 50k / zh 50k):
Column
Type
Description
text
string
Transcript (identical to the original dataset)
audio
Audio
Enhanced audio, FLAC… See the full description on the dataset page: https://huggingface.co/datasets/FILM6912/th-en-zh-tts-200k-enhanced.mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.librispeech_test_clean_enhancedvibravox_enhanced_by_EBEN
Dataset Card
Description
This dataset features a speech-enhanced version of the test split from the speech_clean subset of the Vibravox Dataset.
It is not intended for training.
Enhancement procedure
The Bandwidth extension task has been individually achieved for each sensor using configurable EBEN (arXiv link) models available at https://huggingface.co/Cnam-LMSSC/vibravox_EBEN_models.
Ressources
Results for speech-to-phoneme and speaker… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/vibravox_enhanced_by_EBEN.annotated_catalan_common_voice_v17_cleaned_enhanced
Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of:
projecte-aina/annotated_catalan_common_voice_v17.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.reference-voices-enhanced
Reference Voices Enhanced
2,004 AI voice samples enhanced with ClearerVoice-Studio MossFormer2_SE_48K speech enhancement, annotated with Empathic Insight Voice Plus (59 quality + emotion scores).
Dataset Summary
Source: laion/ai-voices-deduplicated (2,004 speaker-deduplicated, quality-filtered AI voice samples)
Speech Enhancement: ClearerVoice MossFormer2_SE_48K — background noise removal and speech clarity improvement
Output Format: Enhanced WAV files at 48kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/reference-voices-enhanced.voxpopuli_spanish_enhanced
VoxPopuli Spanish Enhanced (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of the Spanish subset of:
facebook/voxpopuli.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been satisfactory.
In addition to… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/voxpopuli_spanish_enhanced.MCE_Mixed_Cantonese_English_Speech_Enhancedcommon_voice_audio_quality_enhancement_v3french-mrt_enhanced_by_EBENSame dataset as Cnam-LMSSC/french-mrt but enhanced by EBEN models
vi-speech-enhancementcommon_voice_audio_quality_enhancementTH-Speech-Enhancedeval-ursa-2-enhanced-eka-hard-20260408-1924
Evaluation Results: ursa-2-enhanced
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
speechmatics/ursa-2-enhanced
34.09%
23.66%
Source Data
Evaluation Dataset: Trelis/eka-hard
Model Evaluated: speechmatics/ursa-2-enhanced
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for this… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-ursa-2-enhanced-eka-hard-20260408-1924.Mara-Jade-Resemble-Enhance-VersionEnhancementDetection_LibriTTS-TestClean_WHAM_TTSjalandhary_asr_enhancedEnhancementDetection_LibriTTS-TestClean_WHAM
Dataset Card for "EnhancementDetection_LibrittsTestCleanWham"
More Information needed
EnhancementDetection_LibriTTS-TestClean_WHAMopenslr_enhancedspeech-enhancementenhanced_facebook_voxpopulik_16k_Whisper_Compatibleeval-ursa-2-enhanced-medical-terms-2025-20260408-1928
Evaluation Results: ursa-2-enhanced
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
speechmatics/ursa-2-enhanced
6.04%
3.29%
Source Data
Evaluation Dataset: Trelis/medical-terms-2025
Model Evaluated: speechmatics/ursa-2-enhanced
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-ursa-2-enhanced-medical-terms-2025-20260408-1928.reuse-enhancedeurospeech-enhanced-dacvae
EuroSpeech parliamentary speech converted to DAC VAE latents
Source
disco-eth/EuroSpeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/laion/eurospeech-enhanced-dacvae.eval-ursa-2-enhanced-multimed-hard-20260408-1933
Evaluation Results: ursa-2-enhanced
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
speechmatics/ursa-2-enhanced
10.46%
6.03%
Source Data
Evaluation Dataset: Trelis/multimed-hard
Model Evaluated: speechmatics/ursa-2-enhanced
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-ursa-2-enhanced-multimed-hard-20260408-1933.enhanced-vocal-burst
