datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speech_400kspeech_dataenglish-casual-speech-sample-south-african-accent
English Casual Speech Sample (South African Accent)
South African crowd-sourced participants respond to questions about their daily lives and activities.
This dataset is a sample of a larger collection from the same data collection campaign.
Changelog
FEB 2026: initial share. ASR (Chirp3) transcripts. WER: 12%
Specs
Speakers: ~550 unique South African speakers
Total duration: ~60 hours
Files sample rate: 48kHz
Actual sample rate: TBD
Language: English (SA… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/english-casual-speech-sample-south-african-accent.samromur_asr
Dataset Card for samromur_asr
Dataset Summary
This is a modfied copy of the dataset from The Language and Voice Laboratory in RU.
This is the first release of the Samrómur Icelandic Speech corpus that contains 100.000 validated utterances.
The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology.
Languages
The audio is in Icelandic.
The… See the full description on the dataset page: https://huggingface.co/datasets/DavidErikMollberg/samromur_asr.speech_40karabic_speech_data_8.tarEuroSpeech-WebDatasetvb_samplescotus-samuel_a_alito_jr-audio
SCOTUS-sim audio: samuel_a_alito_jr
Per-utterance audio clips from Oyez oral-argument mp3s, sliced at
the start_time / stop_time timestamps stored in the companion
scotus-sim/scotus-samuel_a_alito_jr-training dataset.
Alignment
clip_NNNNN.wav in the tarball corresponds exactly to
audio_segments.jsonl[NNNNN] in the training companion dataset.
In metadata.jsonl each row carries the same 0-padded index in idx.
This supersedes the v1 tarball, which had systematic… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-samuel_a_alito_jr-audio.otoSpeech-HQ-full-duplex-samples
Dataset Card for otoSpeech-HQ-full-duplex-samples: Full-Duplex Conversational Speech Dataset Samples
Dataset Summary
otoSpeech-HQ-full-duplex-samples is a curated collection of high-quality full-duplex conversational speech samples designed for commercial and production-oriented use.
This repository is derived from a private subset of otoSpeech and features carefully selected English two-speaker conversations with enhanced audio quality. The samples are intended for… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-HQ-full-duplex-samples.emo_speech_samplelava_datasetzello-public-channels-voice-sample
Zello Public Channels Voice Dataset Sample
Dataset summary
Total audio: 48.23 hours
Total messages: 13207
Breakdown by language
Language
Hours
Messages
Speakers
Channels
ms
10.65
2056
112
7
en
9.54
2990
362
26
id
6.98
1597
157
11
es
6.86
2144
224
21
tl
4.28
1404
78
6
pt
3.92
863
66
11
sw
2.43
182
21
1
ru
1.30
282
42
12
th
0.95
55090
3
zh
0.87
828
312
10
it
0.15
25
19
7
is
0.09
76
25
10
ko
0.05
75
43
13
vi
0.04
42
34
8
fr… See the full description on the dataset page: https://huggingface.co/datasets/zello/zello-public-channels-voice-sample.
