datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-tts-multilingual-emotional-speechSynthdog-Multilingual-100
Synthdog Multilingual
The Synthdog dataset created for training in Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model.
Using the official Synthdog code, we created >1 million training samples for improving OCR capabilities in Large Vision-Language Models.
Dataset Details
We provide the images for download in two .tar.gz files. Download and extract them in folders of the same name (so cat images.tar.gz.* | tar xvzf -C images; tar xvzf… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/Synthdog-Multilingual-100.Multilingual_Speech_Dataset
Multilingual Speech Dataset
Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English
Repository: https://github.com/IS2AI/MultilingualASR
Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.multilingual-librispeech-webdatasetmultilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.multilingual_ocr_erniesdkpiper-plus-multilingual-7lang-v8-dataset
piper-plus multilingual 7-lang v8 training dataset
Preprocessed training dataset for piper-plus zero-shot TTS v8
(342,855 utterances / 3,692 speakers / 7 languages).
data/ holds the full dataset (audio_norm + spec caches, split tar.gz —
concatenate parts then extract). essential/ holds metadata only
(dataset.jsonl + config + CAM++ speaker embeddings + holdout) for fast restore.
Access is gated (manual approval) because the ja subset derives from
MoeSpeech (see… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/piper-plus-multilingual-7lang-v8-dataset.multilingual-in-the-wildmultilingual-in-the-wild-thinkingNayanaDocs-Multilingual-webdatasetMultilingualLibriMix
