datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yt-kids-asr-bench
Child Speech ASR Benchmark
This dataset contains 16 kHz mono FLAC audio embedded in Parquet rows through the Hugging Face Audio feature. It is intended for manually reviewed ASR benchmarking access.
Configs
Config
Split
Rows
Duration
zh
test
18
02:24:00.076
en
test
55
03:43:24.454
v3
test
321
15:43:48.000
v3
zh
319
13:01:23.000
v4
en
55
05:29:50.418
v4
zh
104
05:09:14.000
total
872
45:31:39.948
Columns
audio: embedded… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/yt-kids-asr-bench.CHILDES-Aligned
[!IMPORTANT]
How to access this dataset: the official public release is hosted by TalkBank at
https://talkbank.org/childes/access/Derived/CHILDES-Aligned.html (audio archives +
CSV/JSONL metadata, CC BY-NC-SA 4.0). Please obtain the dataset there.
This Hugging Face copy is retained gated, for internal use; access requests are
approved manually and general requests may be declined — use the TalkBank release instead.
CHILDES-Aligned: Curated Child-Speech Dataset (BEACON)
English… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/CHILDES-Aligned.volga-ipatava-znak-vialikaga-magistra-valiantsin-aksiantsiuk-output_original
Знак Вялікага Магістра — арыгінальнае аўдыё
Аўтар / Author: Вольга ІпатаваМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
volga-ipatava-znak-vialikaga-magistra-valiantsin-aksiantsiuk-output
Доўгасць аўдыё
4h13m
Радкоў у датасеце
1 302
Структура
Кожны радок змяшчае:… See the full description on the dataset page: https://huggingface.co/datasets/fosters/volga-ipatava-znak-vialikaga-magistra-valiantsin-aksiantsiuk-output_original.volga-ipatava-znak-vialikaga-magistra-valiantsin-aksiantsiuk_all
AudioSet Pipeline Output
Мова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
2,373
Працягласць
10 гадз 9 хв
Частата дыскрэтызацыі
22050 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/volga-ipatava-znak-vialikaga-magistra-valiantsin-aksiantsiuk_all.Gleason
TalkBank CHILDES Gleason Raw
This repository mirrors the Gleason corpus from TalkBank CHILDES as raw files
for reproducible local workflows.
Source corpus page: https://talkbank.org/childes/access/Eng-NA/Gleason.html
DOI: doi:10.21415/T5101R
HF repo: MagicLuke/Gleason
Contents
transcripts/Gleason/{Mother,Father,Dinner}/*.cha
media/{Mother,Father,Dinner}/*.mp3
raw/Gleason.zip
metadata.json
metadata_from_cha.json
recordings_from_cha.csv
Data config
The… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/Gleason.Redmond-Sentence-Recall
RSR segmented-latest
This dataset contains segmented child speech utterances built from the
RSR raw-latest CHAT/audio tree with talkbank-toolkit.
Initial upload target: MagicLuke/Redmond-Sentence-Recall (private).
Configs
Config
Split
Rows
Source manifest
sentence
train
14563
train.sentence.jsonl
sentence
test
2065
test.sentence.jsonl
all_chi
train
15102
train.all_chi.jsonl
all_chi
test
2167
test.all_chi.jsonl
Config Meaning… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/Redmond-Sentence-Recall.
