datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
belebele-fleurs
Belebele-Fleurs
Belebele-Fleurs is a dataset suitable to evaluate two core tasks:
Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form.
Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.2M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.bel_canto
Dataset Card for Bel Conto and Chinese Folk Song Singing Tech
Original Content
This dataset is created by the authors and encompasses two distinct singing styles: bel canto and Chinese folk singing. Bel canto is a vocal technique frequently employed in Western classical music and opera, symbolizing the zenith of vocal artistry within the broader Western musical heritage. Chinese folk singing, for which there is no official English translation, is referred to here as a… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/bel_canto.bulgarian-audiobooks-tts-400hBulgarian Female Audiobook TTS Dataset (400 Hours)
Description
This dataset contains over 400 hours of high-quality Bulgarian speech audio, specifically curated for training and fine-tuning Text-to-Speech (TTS) models. The data is sourced from various audiobooks and has undergone a rigorous filtering and cleaning process to ensure high training stability.
Dataset Specifications
Total Duration: ~400 hours (post-filtering)
Total Segments: ~200,000
Segment Length: 4 – 12 seconds
Language:… See the full description on the dataset page: https://huggingface.co/datasets/beleata74/bulgarian-audiobooks-tts-400h.Bulgarian-TTS-1300h
Bulgarian TTS Dataset (1300 Hours)
This is a large-scale, high-quality Bulgarian Text-to-Speech (TTS) dataset containing approximately 1300 hours of transcribed audio.
Dataset Curation & Filtering
The original raw dataset consisted of 3000 hours of audio. To ensure the highest quality for training TTS models (and specifically for optimizing model context windows), a rigorous filtering and cleaning pipeline was applied. The final dataset is 1300 hours, with the… See the full description on the dataset page: https://huggingface.co/datasets/beleata74/Bulgarian-TTS-1300h.aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy
Казкі і апавяданні беларусаў Слуцкага павету
Metadata
Author: Аляксандр Сержпутоўскі
Title: Казкі і апавяданні беларусаў Слуцкага павету
Narrator: Юры Жыгамонт
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy.aaaa111223
VoxAging
VoxAging: Continuously Tracking Speaker Aging with a Large-Scale Longitudinal Dataset in English and Mandarin
📄 Paper: https://arxiv.org/pdf/2505.21445
Description
The VoxAging dataset is a large-scale longitudinal audio-visual corpus designed for studying long-term speaker aging and temporal variations in speech.It consists of recordings from 293 speakers, including 226 English speakers (112 female, 114 male) and 67 Mandarin speakers (23 female, 44… See the full description on the dataset page: https://huggingface.co/datasets/belztjti/aaaa111223.2M-Belebele-Jabelarusian-speech-datasetLouise-Belcher-WAV-Datasetmgb2_audios_transcriptions_prepared
Dataset Card for "mgb2_audios_transcriptions_prepared"
More Information needed
aliaksandr-serzhputouski-prymkhi-i-zababony-belarusau-paleshukou-iury-zhygamont
Прымхі і забабоны беларусаў-палешукоў
Metadata
Author: Аляксандр Сержпутоўскі
Title: Прымхі і забабоны беларусаў-палешукоў
Narrator: Юры Жыгамонт
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-prymkhi-i-zababony-belarusau-paleshukou-iury-zhygamont.bell_ring_put_tape_in_bin_evalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 6817,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AdamMalyshev/bell_ring_put_tape_in_bin_eval.bell_ring_put_tape_in_binThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 30598,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AdamMalyshev/bell_ring_put_tape_in_bin.be-bel-audio-corpusBelarusian audio corpus.
3 different sets under the hood.
dataset_1 - Людзі на балоце
Corresponded text https://knihi.com/Ivan_Mielez/Ludzi_na_balocie.html
Total duration: 13:42:04
dataset_2 - Donar.by
Total duration: 37h 5m 34s
dataset-3 - Knihi.by
Sources (with kind permission by): PRASTORA.BY (https://prastora.by)
Total durations: 36:50:54
eval_bell_ring_put_tape_in_bin_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 19,
"total_frames": 10006,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:19"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AdamMalyshev/eval_bell_ring_put_tape_in_bin_test.hamza-belloumi-tunisian-tts
Hamza Belloumi Tunisian TTS Dataset
A Tunisian Arabic speech dataset for TTS model training.
bely_klyck_all
Белы Клык
Аўтар / Author: Джэк ЛонданМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
2,056
Працягласць
5 гадз 20 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/bely_klyck_all.Alexander-Graham-Bell-1885mgb2_audios_transcriptions_non_overlap
Dataset Card for "mgb2_audios_transcriptions_non_overlap"
More Information needed
bella_ciaomgb2_audios_transcriptions
Dataset Card for "mgb2_audios_transcriptions"
More Information needed
eval_bell_ring_put_tape_in_bin_random_init_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 12,
"total_frames": 6383,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AdamMalyshev/eval_bell_ring_put_tape_in_bin_random_init_test.mgb2_audios_transcriptions_preprocessed
Dataset Card for "mgb2_audios_transcriptions_preprocessed"
More Information needed
bely_klyck_output
Белы Клык
Аўтар / Author: Джэк ЛонданМова / Language: Беларуская (Belarusian)
Частка калекцыі Ministerskija —
выраўнаваныя аўдыёзапісы беларускіх аўдыёкніг з транскрыпцыямі.
Апублікаваных радкоў (HF)
1,863
Агулам у БД
4,831
Доўгасць аўдыё
12h36m
Парог даверу
≥ 0.95
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (~15 с)
text — транскрыпцыя (Gemini + ASR выраўнаванне)
chunk_uid — унікальны ідэнтыфікатар фрагмента
Апрацоўка… See the full description on the dataset page: https://huggingface.co/datasets/fosters/bely_klyck_output.bella_ciaopick_and_place_cup_bell_dropThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 27,
"total_frames": 11196,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"audio_files_size_in_mb": 100,
"fps": 30,
"splits": {
"train": "0:27"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AdamMalyshev/pick_and_place_cup_bell_drop.AI-Belha
Dataset Card for AI-Belha
Dataset Summary
AI-Belha is a dataset comprising audio recordings from beehives, collected to determine the presence and status of the queen bee. The dataset includes 86 mono WAV files, each approximately 60 seconds long and sampled at 16 kHz, totaling about 1 hour and 26 minutes of audio. Each recording is annotated with beekeeper observations and model predictions regarding the queen bee's status.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/NOSInovacao/AI-Belha.ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output
Кароткая гісторыя Беларусі
Аўтар / Author: Алесь КраўцэвічДыктар / Narrator: Уладзімір ЛісоўскіМова / Language: Беларуская (Belarusian)
Частка калекцыі Ministerskija —
выраўнаваныя аўдыёзапісы беларускіх аўдыёкніг з транскрыпцыямі.
Апублікаваных радкоў (HF)
675
Доўгасць аўдыё
2h21m
Парог даверу
≥ 0.95
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (~15 с)
text — транскрыпцыя (Gemini + ASR выраўнаванне)
chunk_uid — унікальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output.ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output_original
Кароткая гісторыя Беларусі — арыгінальнае аўдыё
Аўтар / Author: Алесь КраўцэвічМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output
Доўгасць аўдыё
2h21m
Радкоў у датасеце
675
Структура
Кожны радок… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output_original.
