datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iamd_v0
Internet Archive Music Dataset (IAMD v0)
~4.2M thirty-second music segments (34,469 hours) sourced from
Creative-Commons audio on the Internet Archive, each
paired with machine-generated natural-language captions and the original item
metadata.
Segments
4.2M
Audio
34k hours
Segment length
30 s nominal (mean 29.22 s)
Format
MP3, 320 kbps CBR, native channels + sample rate
Shards
2,320 Parquet files
Download size
4.53 TB
Loading
A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.vzLeViSQA-v1Liv-IA-audiodatasetiany-khmer-voice
iAny Khmer Voice
An open, community-contributed Khmer speech dataset for training speech-to-text
(ASR). Recorded through the iAny "Contribute your voice" page
(https://iany.app/voice), where people read short Khmer sentences aloud —
with the community, for the community.
Clips: 5629
Speakers: 79 (anonymous ids)
Duration: 5.3 hours
Audio: 16 kHz mono WAV
Language: Khmer (km)
License: CC-BY-SA-4.0
Structure
Standard Hugging Face audiofolder:… See the full description on the dataset page: https://huggingface.co/datasets/sengtha/iany-khmer-voice.ru_common_voice_sova_rudevices_golos_fleursvietnamese-music-datasetOpenSLR54-Nepali-ASRIARPA_BABEL_OP3_306MusicAVQA-A2V-Retrieval
MusicAVQA-A2V-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses audio queries and video corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-A2V-Retrieval.Voice_Training_DataA database of the Audio Samples that was use to trained the Voice Models
MusicAVQA-V2A-Retrieval
MusicAVQA-V2A-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses video queries and audio corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-V2A-Retrieval.dataset_ia_strixelise-clone
Custom Elise-like TTS Dataset
Converted on 2025-06-18T05:56:51Z.
Samples: 1043
Format : 10-s clips with text transcription (like MrDragonFox/Elise)
Structure
column
type
description
audio
audio
24kHz mono wav clip
text
string
transcription
LeViSQAViSQA-newiamnotthatkindoftalentiakub-kolas-kazki-zhytstsia-output_original
Казкі жыцця — арыгінальнае аўдыё
Аўтар / Author: Якуб КоласМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
iakub-kolas-kazki-zhytstsia-output
Доўгасць аўдыё
1h49m
Радкоў у датасеце
504
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/iakub-kolas-kazki-zhytstsia-output_original.eightflorapaulitaariDS_STRIX_IAts_datasetIALP-2026-data
IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets
Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC
Adaptation." This repository holds the fixed target-query, validation, and
evaluation sets used across all experiments. Each part is a self-contained
.tar.gz.
All audio is 16 kHz mono. Each split ships with:
audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16)
wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.Cover.IAViSQA_plusfluent_slu_v1.0
Dataset Card for "fluent_slu_v1.0"
More Information needed
24khz_800clipsAffectHuman-43K
AffectHuman-43K
AffectHuman-43K is an emotion-aligned multimodal benchmark for controlled human affect generation and evaluation.
The benchmark contains 42,469 usable samples with complete image, reference-image, audio, and text coverage. Identity is specified through a visual reference image, while text, audio, and emotion labels provide affective control signals. This design separates identity preservation from affective control, enabling evaluation of whether a model can preserve… See the full description on the dataset page: https://huggingface.co/datasets/iamjamuna/AffectHuman-43K.emotional_tts_datasetianka-sipakou-zialeny-listok-na-planetse-ziamlia-maryia-zakharevich
Зялёны лісток на планеце Зямля
Metadata
Author: Янка Сіпакоў
Title: Зялёны лісток на планеце Зямля
Narrator: Марыя Захарэвіч
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/ianka-sipakou-zialeny-listok-na-planetse-ziamlia-maryia-zakharevich.
