child_speech
childvox-speechmaturity-babyhubertchildvox-speechmaturity-whisper-largechildvox-speechmaturity-whisper-basechildvox-speechocean762-accuracy-whisper-largechildvox-speechocean762-prosody-whisper-basechildvox-speechocean762-prosody-babyhubertchildvox-speechocean762-accuracy-babyhubertchildvox-speechocean762-accuracy-whisper-base
Kazakh-Russian-Child-Directed-Speech-Corpus
Kazakh-Russian Child-Directed Language Corpus
Dataset Description
The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization.
The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts.
This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.enni-child-speech-synthesislicense: mit
task_categories:
text-to-speech
automatic-speech-recognition
language:
en
tags:
speech
audio
child-speech
talkbank
size_categories:
10K<n<100K
TalkBank Child Speech Synthesis Dataset (Seed 1)
This dataset contains child speech synthesis data generated from the TalkBank FASA ENNI corpus.
Dataset Information
Number of Samples: 10032
Seed: 1
Audio Format: WAV (16kHz)
Source: TalkBank FASA ENNI
Data Structure
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/jsun39/enni-child-speech-synthesis.child-tts-speechocean-subsetChild_Speech_dataset_Whisper
