datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test321
test321
This is a merged speech dataset containing 118 audio segments from 2 source datasets.
Dataset Information
Total Segments: 118
Speakers: 4
Languages: tr
Emotions: happy, angry, sad, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test321.testpic01
EuroSpeech Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts.
Dataset Summary
Languages: 22 European languages (see detailed breakdown below)
Total aligned hours: ~78,100 hours of… See the full description on the dataset page: https://huggingface.co/datasets/zyzsasasas/testpic01.ts_asr_test
ts_asr_test:目标说话人 ASR 测试集(manifest-only)
ts_asr_test is a 3,928-clip (~8.7 h) Chinese/English target-speaker ASR test set, released
manifest-only: the repo ships no audio, only an audio-free recipe and a self-contained,
deterministic rebuild script. Bring your own copies of the public source corpora and run
rebuild_ts_asr_test.py to regenerate every clip bit-for-bit.
数据集简介
每条样本由一段目标说话人语音、一段同说话人的注册音频(enrollment),以及若干干扰说话人语音与背景噪声按固定配方混合而成。任务:在给定 enrollment… See the full description on the dataset page: https://huggingface.co/datasets/Boxp/ts_asr_test.freddy-testDette er et datasett som skal slettes.
testdataset
NeMo Tarred Dataset
Generated from Test3.
Train rows: 40338 · Test rows: 422
Shards: 5 · Codec: flac · Sample rate: 16000 Hz mono
Primary text: text · target_lang: ta-IN
is_tarred: true
tarred_audio_filepaths: .../audio__OP_0..4_CL_.tar
manifest_filepath: .../train_manifest.json
testtr43
testtr43
This is a merged speech dataset containing 2655 audio segments from 3 source datasets.
Dataset Information
Total Segments: 2655
Speakers: 13
Languages: tr
Emotions: angry, happy, neutral
Original Datasets: 3
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/testtr43.
