datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmnilingualASR-retrieval
Omnilingual ASR speech-text retrieval (MTEB)
Read speech paired with its human transcription, for languages that no existing
MTEB audio task covers.
Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official
test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated
transcripts are dropped, since one would otherwise be relevant to several
recordings while only one is marked correct.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.free-music-archive-retrieval
FMAR: A Dataset for Robust Song Identification
Authors: Ryan Lee, Yi-Chieh Chiu, Abhir Karande, Ayush Goyal, Harrison Pearl, Matthew Hong, Spencer Cobb
Overview
To improve copyright infringement detection, we introduce Free-Music-Archive-Retrieval (FMAR), a structured dataset designed to test a model's capability to identify songs based on 5-second clips, or queries. We create adversarial queries to replicate common strategies to evade copyright infringement detectors… See the full description on the dataset page: https://huggingface.co/datasets/ml-ryanlee/free-music-archive-retrieval.IndicDiarBench-speaker-retrieval
Indic DiarBench speaker retrieval (MTEB)
Indic DiarBench reshaped for speaker retrieval across the 22 scheduled languages
of India: given a clip of one speaker, find other clips of that same speaker.
Source: sarvamai/indic-diarbench at revision 92877ba, cc-by-4.0, official
test split. Turns are cut by their annotated times, restricted to 2 to 15
seconds, and turns overlapping a different speaker are dropped. identity pairs
the recording session with the speaker, because speaker… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/IndicDiarBench-speaker-retrieval.LinguaLibre-word-retrieval
Lingua Libre spoken-word retrieval (MTEB)
Single words read aloud by volunteers, paired with the written word, across many
languages.
Recordings come from Lingua Libre, a Wikimedia project, hosted on Wikimedia
Commons, which is free by site policy. Published as cc-by-sa-4.0. Audio is 16 kHz
Opus. Bare punctuation and read sentences are excluded, and each word is kept once.
Built by scripts/data/lingua_libre/create_data.py in the MTEB repo.
AfriMCQA-speech-image-retrieval
Afri-MCQA speech-image retrieval (MTEB)
Afri-MCQA reshaped for retrieval: find the photograph a spoken question is asking
about. Questions are spoken by native speakers in 16 African languages and are
grounded in culturally relevant images.
Source: Atnafu/Afri-MCQA at revision 8b8c53d, cc-by-nc-4.0, official
test split. Images are stored per language because one photograph can carry
questions in several languages.
Built by scripts/data/afri_mcqa_retrieval/create_data.py in the… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-speech-image-retrieval.vaani-audio-image-retrieval
Vaani audio–image retrieval (MTEB)
Multilingual audio↔image retrieval over 62 Indian languages, derived from
Project Vaani (IISc Bangalore /
ARTPARK).
Vaani records image-prompted speech: a speaker is shown a photograph and describes it
aloud in their own language. Each recording is therefore grounded in a specific image,
which is what makes audio↔image retrieval well defined without any extra annotation.
Prepared for MTEB as
VaaniA2IRetrieval and VaaniI2ARetrieval.… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vaani-audio-image-retrieval.vaani-speech-text-retrievalavcaps-retrieval
AVCaps audio–visual retrieval (MTEB)
Retrieval tasks over AVCaps, an
audio-visual dataset derived from VidOR in which each clip is captioned three separate
ways — from the audio alone, from the visuals alone, and from both together.
That separation is the point: the audio-only, video-only and combined directions can be
scored independently on identical clips, rather than inferred from a single caption set
that mixes the modalities.
Prepared for MTEB as six tasks:… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/avcaps-retrieval.waxal-audio-text-retrieval
WAXAL speech–text retrieval (MTEB)
Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from
WAXAL (Google and partners).
Most of these languages have no presence in mteb's existing multilingual audio tasks,
which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and
WaxalT2ARetrieval.
Contents
One config per language, each with id, audio, text, speaker_id, gender,
language. 1,722 utterances total.
code
language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.VGGSound_AV_RETRIEVALSpokenWikipedia-retrieval
Spoken Wikipedia speech-text retrieval (MTEB)
Volunteer readings of Wikipedia articles paired with the article lead, in Dutch,
English, German, Spanish and French.
Recordings come from Wikimedia Commons, which is free by site policy, and the lead
text from each Wikipedia, which is CC-BY-SA. The set is published as cc-by-sa-4.0.
Only the first 60 seconds of each reading is kept, since readers start
at the lead. One recording per article.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/SpokenWikipedia-retrieval.LombardGrid-Retrieval
Lombard GRID Cross-Modal Utterance Retrieval
Four MTEB/MOEB utterance-retrieval tasks derived from the
Lombard GRID corpus: audio-to-video, video-to-audio,
frontal-to-profile video, and plain-to-Lombard video+audio. Relevance always
requires the same utterance recording or the same speaker-and-sentence pair;
speaker identity alone is never relevant.
The source paper introduced the corpus but did not define these retrieval tasks.
Relevance is derived from native recording… See the full description on the dataset page: https://huggingface.co/datasets/Cerru02/LombardGrid-Retrieval.MusicAVQA-A2V-Retrieval
MusicAVQA-A2V-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses audio queries and video corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-A2V-Retrieval.MCIF-retrieval
MCIF multilingual audio-visual retrieval (MTEB)
MCIF reshaped for retrieval: find the recorded conference-talk segment that
answers a question asked in English, German, Italian or Chinese. The questions
are parallel across the four languages while the talks are spoken in English,
so the non-English subsets measure cross-lingual grounding.
Source: FBK-MT/MCIF at revision e24065b, cc-by-4.0. Only the
question-answering samples are used, restricted to answerable and… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/MCIF-retrieval.MusicAVQA-V2A-Retrieval
MusicAVQA-V2A-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses video queries and audio corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-V2A-Retrieval.lass-synth-retrieval-mini
