CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vnahata /OmnilingualASR-retrieval Omnilingual ASR speech-text retrieval (MTEB) Read speech paired with its human transcription, for languages that no existing MTEB audio task covers. Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated transcripts are dropped, since one would otherwise be relevant to several recordings while only one is marked correct. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.audioautomatic-speech-recognition1K<n<10K0 likes1.1k downloads22d agoHugging Face02ml-ryanlee /free-music-archive-retrieval FMAR: A Dataset for Robust Song Identification Authors: Ryan Lee, Yi-Chieh Chiu, Abhir Karande, Ayush Goyal, Harrison Pearl, Matthew Hong, Spencer Cobb Overview To improve copyright infringement detection, we introduce Free-Music-Archive-Retrieval (FMAR), a structured dataset designed to test a model's capability to identify songs based on 5-second clips, or queries. We create adversarial queries to replicate common strategies to evade copyright infringement detectors… See the full description on the dataset page: https://huggingface.co/datasets/ml-ryanlee/free-music-archive-retrieval.audioaudio-classification1K<n<10K1 likes568 downloads1y agoHugging Face03vnahata /IndicDiarBench-speaker-retrieval Indic DiarBench speaker retrieval (MTEB) Indic DiarBench reshaped for speaker retrieval across the 22 scheduled languages of India: given a clip of one speaker, find other clips of that same speaker. Source: sarvamai/indic-diarbench at revision 92877ba, cc-by-4.0, official test split. Turns are cut by their annotated times, restricted to 2 to 15 seconds, and turns overlapping a different speaker are dropped. identity pairs the recording session with the speaker, because speaker… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/IndicDiarBench-speaker-retrieval.audioaudio-classification1K<n<10K0 likes509 downloads22d agoHugging Face04vnahata /LinguaLibre-word-retrieval Lingua Libre spoken-word retrieval (MTEB) Single words read aloud by volunteers, paired with the written word, across many languages. Recordings come from Lingua Libre, a Wikimedia project, hosted on Wikimedia Commons, which is free by site policy. Published as cc-by-sa-4.0. Audio is 16 kHz Opus. Bare punctuation and read sentences are excluded, and each word is kept once. Built by scripts/data/lingua_libre/create_data.py in the MTEB repo. audioautomatic-speech-recognition1K<n<10K1 likes424 downloads21d agoHugging Face05vnahata /AfriMCQA-speech-image-retrieval Afri-MCQA speech-image retrieval (MTEB) Afri-MCQA reshaped for retrieval: find the photograph a spoken question is asking about. Questions are spoken by native speakers in 16 African languages and are grounded in culturally relevant images. Source: Atnafu/Afri-MCQA at revision 8b8c53d, cc-by-nc-4.0, official test split. Images are stored per language because one photograph can carry questions in several languages. Built by scripts/data/afri_mcqa_retrieval/create_data.py in the… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-speech-image-retrieval.audioimage-to-text1K<n<10K0 likes416 downloads22d agoHugging Face06vnahata /vaani-audio-image-retrieval Vaani audio–image retrieval (MTEB) Multilingual audio↔image retrieval over 62 Indian languages, derived from Project Vaani (IISc Bangalore / ARTPARK). Vaani records image-prompted speech: a speaker is shown a photograph and describes it aloud in their own language. Each recording is therefore grounded in a specific image, which is what makes audio↔image retrieval well defined without any extra annotation. Prepared for MTEB as VaaniA2IRetrieval and VaaniI2ARetrieval.… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vaani-audio-image-retrieval.audioaudio-classification1K<n<10K0 likes338 downloads24d agoHugging Face07vnahata /vaani-speech-text-retrievalaudio1K<n<10K0 likes257 downloads23d agoHugging Face08vnahata /avcaps-retrieval AVCaps audio–visual retrieval (MTEB) Retrieval tasks over AVCaps, an audio-visual dataset derived from VidOR in which each clip is captioned three separate ways — from the audio alone, from the visuals alone, and from both together. That separation is the point: the audio-only, video-only and combined directions can be scored independently on identical clips, rather than inferred from a single caption set that mixes the modalities. Prepared for MTEB as six tasks:… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/avcaps-retrieval.audiotext-to-audio1K<n<10K0 likes161 downloads24d agoHugging Face09vnahata /waxal-audio-text-retrieval WAXAL speech–text retrieval (MTEB) Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from WAXAL (Google and partners). Most of these languages have no presence in mteb's existing multilingual audio tasks, which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and WaxalT2ARetrieval. Contents One config per language, each with id, audio, text, speaker_id, gender, language. 1,722 utterances total. code language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.audioautomatic-speech-recognition1K<n<10K0 likes131 downloads24d agoHugging Face10mteb /VGGSound_AV_RETRIEVALaudion<1K0 likes119 downloads7mo agoHugging Face11vnahata /SpokenWikipedia-retrieval Spoken Wikipedia speech-text retrieval (MTEB) Volunteer readings of Wikipedia articles paired with the article lead, in Dutch, English, German, Spanish and French. Recordings come from Wikimedia Commons, which is free by site policy, and the lead text from each Wikipedia, which is CC-BY-SA. The set is published as cc-by-sa-4.0. Only the first 60 seconds of each reading is kept, since readers start at the lead. One recording per article. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/SpokenWikipedia-retrieval.audioautomatic-speech-recognitionn<1K0 likes108 downloads21d agoHugging Face12Cerru02 /LombardGrid-Retrieval Lombard GRID Cross-Modal Utterance Retrieval Four MTEB/MOEB utterance-retrieval tasks derived from the Lombard GRID corpus: audio-to-video, video-to-audio, frontal-to-profile video, and plain-to-Lombard video+audio. Relevance always requires the same utterance recording or the same speaker-and-sentence pair; speaker identity alone is never relevant. The source paper introduced the corpus but did not define these retrieval tasks. Relevance is derived from native recording… See the full description on the dataset page: https://huggingface.co/datasets/Cerru02/LombardGrid-Retrieval.audioany-to-any1K<n<10K0 likes87 downloads21d agoHugging Face13iamfortytwo /MusicAVQA-A2V-Retrieval MusicAVQA-A2V-Retrieval This is a derived retrieval benchmark from the test split of mteb/MUSIC-AVQA_cls-preprocessed at revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses audio queries and video corpus items. Construction The source clips are labelled with 22 musical-instrument classes. For every class, a deterministic seed (42) selects five clips as queries and ten distinct clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-A2V-Retrieval.audio1K<n<10K0 likes69 downloads9d agoHugging Face14vnahata /MCIF-retrieval MCIF multilingual audio-visual retrieval (MTEB) MCIF reshaped for retrieval: find the recorded conference-talk segment that answers a question asked in English, German, Italian or Chinese. The questions are parallel across the four languages while the talks are spoken in English, so the non-English subsets measure cross-lingual grounding. Source: FBK-MT/MCIF at revision e24065b, cc-by-4.0. Only the question-answering samples are used, restricted to answerable and… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/MCIF-retrieval.audiotext-to-videon<1K0 likes64 downloads22d agoHugging Face15iamfortytwo /MusicAVQA-V2A-Retrieval MusicAVQA-V2A-Retrieval This is a derived retrieval benchmark from the test split of mteb/MUSIC-AVQA_cls-preprocessed at revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses video queries and audio corpus items. Construction The source clips are labelled with 22 musical-instrument classes. For every class, a deterministic seed (42) selects five clips as queries and ten distinct clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-V2A-Retrieval.audio1K<n<10K0 likes47 downloads9d agoHugging Face16diffunity /lass-synth-retrieval-miniaudio1K<n<10K0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.