datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
symile-m3
Dataset Card for Symile-M3
Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice.
Paper: https://arxiv.org/abs/2411.01053
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.GTSinger
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University
Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.AISHELL-3AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be used to train multi-speaker Text-to-Speech (TTS) systems.The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers and total 88035 utterances. Their auxiliary attributes such as gender, age group and native accents are explicitly marked and provided in the corpus. Accordingly, transcripts in… See the full description on the dataset page: https://huggingface.co/datasets/AISHELL/AISHELL-3.LAION-Audio-300Mgenshin-voice
Genshin Voice
Genshin Voice is a dataset of voice lines from the popular game Genshin Impact.
Hugging Face 🤗 Genshin-Voice
ModelScope Genshin-Voice
Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index.
Last update at 2026-08-13
654252 wavs
7291 without speaker (1%)
52693 without transcription (8%)
1088 without inGameFilename (0%)
Dataset Details
Dataset Description
The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.coral-v3
CoRal: Danish Conversational and Read-aloud Dataset
Version 3.0
Dataset Overview
CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v3.esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.audio-filesshrutilipizenless-voice
Zenless Voice
Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero.
Hugging Face 🤗 Zenless-Voice
ModelScope Zenless-Voice
Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index.
Last update at 2026-09-17, game version 3.2.0
406720 wavs
78785 without speaker (19%)
123429 without transcription (30%)
83509 without inGameFilename (21%)
Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.MRSDrama
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
Yu Zhang*, Wenxiang Guo*, Changhao Pan*, Zhiyuan Zhu*, Tao Jin, Zhou Zhao | Zhejiang University
Dataset of ISDrama (ACMMM 2025): Immersive Spatial Drama Generation through Multimodal Prompting.
We construct MRSDrama, the first multimodal recorded spatial drama dataset, containing binaural drama audios, scripts, videos, geometric poses, and textual prompts.
We provide the full corpus… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/MRSDrama.genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.AISHELL-3
AISHELL-3
Identifier: SLR93
Summary: Mandarin data, provided by Beijing Shell Shell Technology Co., Ltd.
Category: Speech
License: Apache License v.2.0
Downloads (use a mirror closer to you):
data_aishell3.tgz [19G] (speech data and transcripts
) Mirrors:
[US]
[EU]
[CN]
About this resource:AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus
published by Beijing Shell Shell Technology Co.,Ltd. It can be… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-3.numberblocks-one-voice-datasetmultilingual-speech-commands-3lang-raw
Multilingual Speech Commands Dataset (3 Languages, Raw)
This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied.
All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.dnr-v3-multilingualliepa-3
LIEPA-3 — Lithuanian Speech Corpus
Didysis lietuvių kalbos garsynas (LIEPA-3)
Dataset Summary
LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours,
~7.5 million audio files) built for automatic speech recognition (ASR),
text-to-speech (TTS) and linguistic research. It spans read, spontaneous,
phonetically-annotated and dialectal speech recorded under a wide range of
conditions (studio, dictaphone, radio, TV, telephone, audiobooks).
Official… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-3.Neko_Audio-30K_Longmuaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']
وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.starrail-voice
StarRail Voice
StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail.
Hugging Face 🤗 StarRail-Voice
ModelScope StarRail-Voice
Last update at 2026-07-16, game version 4.4.0
403437 wavs
60164 without speaker (15%)
61375 without transcription (15%)
57869 without inGameFilename (14%)
Dataset Details
Dataset Description
The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.ISSAI_KSC_335RS_v_1_1
Dataset Card for "ISSAI_KSC_335RS_v_1_1"
Kazakh Speech Corpus (KSC)
Identifier: SLR102
Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours)
Category: Speech
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN]
About this resource:
A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.mp3MMEB-V3
MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models
🌐 Website |
GitHub |
🏆 Leaderboard |
📖 MMEB-V3 Paper |
📖 MMEB-V2 Paper |
📖 MMEB-V1 Paper |
🤗 Models
Introduction
MMEB-V3 is a comprehensive benchmark for evaluating omni-modality embedding models across text, image, video, audio, visual-document, and agent-centric retrieval scenarios.
Building upon MMEB-V1 and MMEB-V2, MMEB-V3 adds 111 new tasks, resulting in 190 evaluation tasks in… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MMEB-V3.warsh-segments-v3
Haitam03/warsh-v3
Warsh (Rewayat Warsh A'n Nafi') Quran recitation, segmented at waqf with
obadx/recitation-segmenter-v2.
Built with warsh-data.
Layout
path
what
data/<reciter>/<surah>.parquet
one file per source recording, audio embedded as 16 kHz mono FLAC
raw/<reciter>/<surah>.mp3
the source recording it came from
segment_params.json
the settings this corpus was produced with
One parquet per source recording, named after it, so re-running a… See the full description on the dataset page: https://huggingface.co/datasets/Haitam03/warsh-segments-v3.GLOBE_V3
Important notice
Differences between V3 version and two previous versions (V1|V2):
This version is built base on Common Voice 21.0 English Subset.
This version only includes utterance that are an exact match with the transcription from Whisper V3 LARGE (CER == 0).
This version includes the original Common Voice metadata (age, gender, accent, and ID).
All audio files in this version are at 24kHz sampling rate.
All audio files in this version are unenhanced. (We’d greatly… See the full description on the dataset page: https://huggingface.co/datasets/MushanW/GLOBE_V3.Music-POSTPROCESS-32eadf7esmart-turn-data-v3.2-trainTraining dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train.emo_webds_2emo_parler
