24khz
Datasets
All datasets matching “24khz”EuroSpeech-24kHz
EuroSpeech 24 kHz Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
Dataset Summary
Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.latam-spanish-speech-orpheus-tts-24khz
LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready)
Dataset Description
This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate.
The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.encodec_24khz-opt-125m-pretrained-ft-librispeech_asr
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr"
More Information needed
lahgtna-arabic-tts-24khz
Lahgtna Arabic TTS — cleaned, 24 kHz
A TTS-ready filtering of oddadmix/dialectal-arabic-lahgtna-v2,
prepared for finetuning Qwen/Qwen3-TTS-12Hz-0.6B-Base
on Arabic dialects.
162,641 utterances · 591.4 hours · 13 dialects · 24 kHz mono — the survivors
of a nine-stage cascade applied to the full 608,121-utterance / 2,934-hour
source corpus. Overall yield: 26.7%.
The source is an ASR corpus. ASR models learn to ignore noise, reverb and
overlapping speech; TTS models learn to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/lahgtna-arabic-tts-24khz.encodec_24khz-librispeech_asr-train.clean.100-features
Dataset Card for "encodec_24khz-librispeech_asr-train.clean.100-features"
More Information needed
malayalam-orpheus-24khz
