datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vai-speech-text-parallel
Vai Speech-Text Parallel Dataset
Dataset Description
This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Vai - vai
Task: Speech Recognition, Text-to-Speech
Size: 23286 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/vai-speech-text-parallel.vais1000
unofficial mirror of VAIS-1000
official announcement: https://vais.vn/vi/tai-ve/hts_for_vietnamese (dead)
mirror: https://github.com/undertheseanlp/text_to_speech/tree/run/data/vais1000/raw
small only 1h40min audio - 1 speaker (female northern accent) - 1k samples
pre-process: none
need to do: check misspelling, restore foreign words phonetised to vietnamese
usage with HuggingFace:
# pip install -q "datasets[audio]"
from datasets import load_dataset
from torch.utils.data import… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vais1000.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.story_audio_4hr_15_diavaani-bihar_vaishali-cleanedstory_audio_4hr_15secdia_american_male_10toynowsarvam-tts-dataset
Sarvam TTS Training Dataset
High-quality TTS training dataset built for expressive speech synthesis.
Stats
Total: 377 segments | 167.9 minutes
English (en-IN): 191 segments | 84.3 minutes
Hindi (hi-IN): 186 segments | 83.6 minutes
Rejection rate: 26.9% after full manual human review of all 483 segments
Emotion Distribution
neutral: 259 | calm: 36 | sad: 32 | angry: 23 | happy: 12 | surprised: 6 | fearful: 6 | excited: 3
How it was… See the full description on the dataset page: https://huggingface.co/datasets/Vaidik7781/sarvam-tts-dataset.mikhas-lynkou-pra-smelaga-vaiaku-mishku-i-iago-slaunykh-tavaryshau
Пра смелага ваяку Мішку і яго слаўных таварышаў
Metadata
Author: Міхась Лынькоў
Title: Пра смелага ваяку Мішку і яго слаўных таварышаў
Narrator:
Source Group: Дзіцячыя
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/mikhas-lynkou-pra-smelaga-vaiaku-mishku-i-iago-slaunykh-tavaryshau.indic_hindi_new_2vaistory_audio_4hr_15_dia_10secdia_british_female_10vaiNeymarjaishankardia_american_female_10Baby-Cry-Classification-Baidumaksim-garetski-na-imperyalistychnai-vaine-maksim-viniarski
На імперыалістычнай вайне
Metadata
Author: Максім Гарэцкі
Title: На імперыалістычнай вайне
Narrator: Максім Вінярскі
Source Group: Аўдыёкнігі
Source: rutracker.org
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/maksim-garetski-na-imperyalistychnai-vaine-maksim-viniarski.ivan-navumenka-khloptsy-samai-vialikai-vainy-andrei-kaliada
Хлопцы самай вялікай вайны
Metadata
Author: Іван Навуменка
Title: Хлопцы самай вялікай вайны
Narrator: Андрэй Каляда
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/ivan-navumenka-khloptsy-samai-vialikai-vainy-andrei-kaliada.my_audio_datasetindic_hindiindic_hindi_finaldia_british_10vais1000_sidon_noise_removalvaishnawi_tts_cleaned_audioeva-vezhnavets-pa-shto-idzesh-voucha-zui-vaitsiakhouskaia
Па што ідзеш, воўча
Metadata
Author: Ева Вежнавец
Title: Па што ідзеш, воўча
Narrator: Зуй-Вайцяхоўская
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/eva-vezhnavets-pa-shto-idzesh-voucha-zui-vaitsiakhouskaia.iuia-vislander-tumas-vislander-zmitser-vaitsiushkevich
Тумас Вісландэр
Metadata
Author: Юя Вісландэр
Title: Тумас Вісландэр
Narrator: Зміцер Вайцюшкевіч
Source Group: Дзіцячыя
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size:… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/iuia-vislander-tumas-vislander-zmitser-vaitsiushkevich.mikhas-lynkou-pra-smelaga-vaiaku-mishku
Пра сьмелага ваяку Мішку
Metadata
Author: Міхась Лынькоў
Title: Пра сьмелага ваяку Мішку
Narrator:
Source Group: Дзіцячыя
Source: http://staroeradio.ru
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/mikhas-lynkou-pra-smelaga-vaiaku-mishku.aliaksandar-vaitovich-mae-akademii
Мае Акадэміі
Metadata
Author: Аляксандар Вайтовіч
Title: Мае Акадэміі
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandar-vaitovich-mae-akademii.
