datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-tts-multilingual-emotional-speechmulti_round_speech_180kMultilingual_Speech_Dataset
Multilingual Speech Dataset
Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English
Repository: https://github.com/IS2AI/MultilingualASR
Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.multi_accent_speech
Multi-Accent English Speech Corpus (Augmented & Speaker-Disjoint)
This dataset is a curated and augmented multi-accent English speech corpus designed for speech recognition, accent classification, and representation learning.It consolidates multiple open-source accent corpora, converts all audio to a unified format, applies targeted data augmentation, and exports in a tidy, Hugging Face–ready structure.
✨ Key Features
Accents covered (12 total):american_english… See the full description on the dataset page: https://huggingface.co/datasets/cagatayn/multi_accent_speech.multilingual-librispeech-webdatasetmultilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.MultiFoley-VGGSound-Test-Audio
Video-Guided Foley Sound Generation with Multimodal Controls
Paper & Project page
This dataset contains the generated results of our MultiFoley work on the filtered VGGSound test cases. We generate 4 samples for each 8s video (we use the first 8s video for evaluation).
The results are generated with both silent video inputs and text inputs (we use the VGGSound category name for simplicity).
Each wave file is named in the format of {category_name}/{u_id}_{start_time}_{idx}.wav, where… See the full description on the dataset page: https://huggingface.co/datasets/czyang/MultiFoley-VGGSound-Test-Audio.multilingual-in-the-wildpaskal.audio.multispeakermultilingual-in-the-wild-thinkingMultilingualLibriMix
