datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmnilingualASR-retrieval
Omnilingual ASR speech-text retrieval (MTEB)
Read speech paired with its human transcription, for languages that no existing
MTEB audio task covers.
Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official
test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated
transcripts are dropped, since one would otherwise be relevant to several
recordings while only one is marked correct.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.LinguaLibre-word-retrieval
Lingua Libre spoken-word retrieval (MTEB)
Single words read aloud by volunteers, paired with the written word, across many
languages.
Recordings come from Lingua Libre, a Wikimedia project, hosted on Wikimedia
Commons, which is free by site policy. Published as cc-by-sa-4.0. Audio is 16 kHz
Opus. Bare punctuation and read sentences are excluded, and each word is kept once.
Built by scripts/data/lingua_libre/create_data.py in the MTEB repo.
waxal-audio-text-retrieval
WAXAL speech–text retrieval (MTEB)
Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from
WAXAL (Google and partners).
Most of these languages have no presence in mteb's existing multilingual audio tasks,
which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and
WaxalT2ARetrieval.
Contents
One config per language, each with id, audio, text, speaker_id, gender,
language. 1,722 utterances total.
code
language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.SpokenWikipedia-retrieval
Spoken Wikipedia speech-text retrieval (MTEB)
Volunteer readings of Wikipedia articles paired with the article lead, in Dutch,
English, German, Spanish and French.
Recordings come from Wikimedia Commons, which is free by site policy, and the lead
text from each Wikipedia, which is CC-BY-SA. The set is published as cc-by-sa-4.0.
Only the first 60 seconds of each reading is kept, since readers start
at the lead. One recording per article.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/SpokenWikipedia-retrieval.
