datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EuroSpeech-24kHz
EuroSpeech 24 kHz Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
Dataset Summary
Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.latam-spanish-speech-orpheus-tts-24khz
LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready)
Dataset Description
This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate.
The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.encodec_24khz-opt-125m-pretrained-ft-librispeech_asr
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr"
More Information needed
lahgtna-arabic-tts-24khz
Lahgtna Arabic TTS — cleaned, 24 kHz
A TTS-ready filtering of oddadmix/dialectal-arabic-lahgtna-v2,
prepared for finetuning Qwen/Qwen3-TTS-12Hz-0.6B-Base
on Arabic dialects.
162,641 utterances · 591.4 hours · 13 dialects · 24 kHz mono — the survivors
of a nine-stage cascade applied to the full 608,121-utterance / 2,934-hour
source corpus. Overall yield: 26.7%.
The source is an ASR corpus. ASR models learn to ignore noise, reverb and
overlapping speech; TTS models learn to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/lahgtna-arabic-tts-24khz.encodec_24khz-librispeech_asr-train.clean.100-features
Dataset Card for "encodec_24khz-librispeech_asr-train.clean.100-features"
More Information needed
malayalam-orpheus-24khzTV-24kHz-2025.12-Neutral-FT-Mini
Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini
Overview
This dataset is a small, high-quality fine-tuning dataset created specifically for speaker refinement and voice matching in Orpheus TTS models.
It consists of 60 newly recorded German speech samples, spoken in a neutral, relaxed, everyday style, closely reflecting the natural speaking voice of the original speaker.
This dataset is intended for:
Speaker adaptation and voice refinement
Fine-tuning Orpheus TTS models… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini.encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features"
More Information needed
iapp-Thai-EN_SpeechDataset_594K_snac_24khz_tokenisedTV-24kHz-Neutral
Thorsten-Voice TV-24kHz-Neutral Dataset
This dataset is a resampled version of the "TV-2022.10-Neutral" configuration from the original Thorsten-Voice TV-44kHz-Full dataset, converted from 44.1kHz to 24kHz sampling rate.
Dataset Description
The Thorsten-Voice dataset contains German speech recordings by Thorsten Müller, suitable for text-to-speech (TTS) training and other speech synthesis tasks.
Changes from Original
Sample Rate: Converted from… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-Neutral.IndicTTS-Hindi-24kHzspanish_tts_noauddataset_24khzukrainian-tts-audiobooks-24khz
Ukrainian Audiobook TTS Dataset (24 kHz)
Description
Ukrainian speech dataset for TTS and ASR tasks.
Source Dataset
https://huggingface.co/datasets/Yehor/audiobooks-xxl
Processing Pipeline
MusicDetection filtering — removed samples with background music/noise
Audio processing (Sidon) — resampled 16 kHz → 24 kHz, converted to mono
Transcription — generated with nvidia/canary-1b-v2
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ukrainian-tts-audiobooks-24khz.encodec_24khz-librispeech_asr100h
Dataset Card for "encodec_24khz-librispeech_asr100h"
More Information needed
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features"
More Information needed
expressive_speech_24khzBernd-Ungerer-24khz14 hours of audio data randomly selected from the data of the speaker "Bernd Ungerer"
original dataset
nurc_tts_24khzencodec_24khz-b24.0-librispeech_asr-features
Dataset Card for "encodec_24khz-b24.0-librispeech_asr-features"
More Information needed
hausa-tts-24khz-waxalnlp-5
vaghawan/hausa-tts-24khz-waxalnlp-5
Single-speaker Hausa TTS dataset from WaxalNLP (google/WaxalNLP hau_tts), speaker 5.
Audio is stored as 24000 Hz FLAC.
Splits
train: 251 rows
hausa-tts-24khz-waxalnlp-3-clean
vaghawan/hausa-tts-24khz-waxalnlp-3-clean
Single-speaker Hausa TTS dataset from WaxalNLP (google/WaxalNLP hau_tts), speaker 3.
Audio is stored as 24000 Hz FLAC.
Splits
train: 441 rows
hausa-tts-24khz-naijavoices-O0456
vaghawan/hausa-tts-24khz-naijavoices-O0456
Single-speaker Hausa TTS dataset from NaijaVoices, speaker O0456.
Audio is stored as 24000 Hz FLAC.
Splits
train: 7682 rows
TV-24kHz-Neutral-tokenised
Thorsten-Voice TV-24kHz-Neutral-tokenised
Overview
This dataset is a tokenised German text-to-speech dataset created for training and fine-tuning the Orpheus TTS model family.
It is based on approximately 12,000 speech recordings from the original Thorsten-Voice Dataset (2022.10) and has been resampled to 24 kHz and tokenised using Orpheus TTS preprocessing.
This dataset is intended for:
Training and fine-tuning Orpheus-based German TTS models
Research on neural speech… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-Neutral-tokenised.hausa-tts-24khz-waxalnlp-3
vaghawan/hausa-tts-24khz-waxalnlp-3
Single-speaker Hausa TTS dataset from WaxalNLP (google/WaxalNLP hau_tts), speaker 3.
Audio is stored as 24000 Hz FLAC.
Splits
train: 246 rows
Eva-K-24khz13 hours of audio data randomly selected from the data of the speaker "Eva K"
original dataset
encodec_24khz-librispeech_asr-validation.clean-features
Dataset Card for "encodec_24khz-librispeech_asr-validation.clean-features"
More Information needed
spanish_tts_dataset_24khzencodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features"
More Information needed
beyoru_kafka-voice-en-24khz-emotesThis dataset is a modified version of beyoru/kafka-voice-en, enhanced with additional features and preprocessing for improved usability in speech-related machine learning tasks.
Modifications
The original dataset has been modified as follows:
Re-transcribed with emotive annotations (e.g., <laughs>, <sighs>)
Audio converted to 24kHz sample rate
Unique identifier (UUID) added for each sample
Audio Evaluation Scores (AES) calculated using Meta's AudioBox Aesthetics project
Short… See the full description on the dataset page: https://huggingface.co/datasets/Gapeleon/beyoru_kafka-voice-en-24khz-emotes.TV-24kHz-2025.12-Neutral-FT-Mini-tokenised
Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini-tokenised
Overview
This dataset is the tokenised version of the TV-24kHz-2025.12-Neutral-FT-Mini dataset, prepared specifically for direct fine-tuning of Orpheus TTS Thorsten-Voice.
It contains 60 tokenised German speech samples, optimised for fast experimentation and precise speaker adaptation.
This dataset is intended for:
Lightweight Orpheus TTS fine-tuning
Speaker identity refinement
Prosody and articulation adjustments… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini-tokenised.
