datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iamd_v0
Internet Archive Music Dataset (IAMD v0)
~4.2M thirty-second music segments (34,469 hours) sourced from
Creative-Commons audio on the Internet Archive, each
paired with machine-generated natural-language captions and the original item
metadata.
Segments
4.2M
Audio
34k hours
Segment length
30 s nominal (mean 29.22 s)
Format
MP3, 320 kbps CBR, native channels + sample rate
Shards
2,320 Parquet files
Download size
4.53 TB
Loading
A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.LeViSQA-v1MusicAVQA-A2V-Retrieval
MusicAVQA-A2V-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses audio queries and video corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-A2V-Retrieval.OpenSLR54-Nepali-ASRMusicAVQA-V2A-Retrieval
MusicAVQA-V2A-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses video queries and audio corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-V2A-Retrieval.elise-clone
Custom Elise-like TTS Dataset
Converted on 2025-06-18T05:56:51Z.
Samples: 1043
Format : 10-s clips with text transcription (like MrDragonFox/Elise)
Structure
column
type
description
audio
audio
24kHz mono wav clip
text
string
transcription
LeViSQAiamnotthatkindoftalentts_datasetViSQA-newViSQA_pluseightfluent_slu_v1.0
Dataset Card for "fluent_slu_v1.0"
More Information needed
emotional_tts_datasetEsp_datasetAffectHuman-43K
AffectHuman-43K
AffectHuman-43K is an emotion-aligned multimodal benchmark for controlled human affect generation and evaluation.
The benchmark contains 42,469 usable samples with complete image, reference-image, audio, and text coverage. Identity is specified through a visual reference image, while text, audio, and emotion labels provide affective control signals. This design separates identity preservation from affective control, enabling evaluation of whether a model can preserve… See the full description on the dataset page: https://huggingface.co/datasets/iamjamuna/AffectHuman-43K.elise-modifiedvietnamese-music-datasetiammusicLaila_low_pausevanessaLaila_600ms_medium_pause24khz_800clipsanna11nepali_to_english_pipeline_evaluation
Nepali-English Speech-to-Text Translation Evaluation Dataset
Dataset Description
This dataset is designed for evaluating Nepali→English speech-to-text translation pipelines.
It contains audio recordings of 300 Nepali sentences, spoken by three speakers, covering a range of sentence types (statements, questions, commands, complex sentences, and named entities/numbers).
Each sentence is paired with:
Source text (Nepali) transcription
Reference English translation
Audio… See the full description on the dataset page: https://huggingface.co/datasets/iamTangsang/nepali_to_english_pipeline_evaluation.Laila_audio_formatedlaila_80_testannadatav2.1IA_MORAES
