CoolFace
Datasetpublic

rakshi719/SpokenCOCO-A2IT

SpokenCOCO Audio-to-(Image+Text) Retrieval MTEB/MOEB task where queries are spoken audio captions and corpus items contain both a MSCOCO image and its written text caption. Task Given a spoken audio description of an image, retrieve the correct (image, text) pair from the corpus. Only models that can process all three modalities — audio, image, and text — can exploit the full corpus signal. Contents Queries: 25031 spoken audio captions (WAV… See the full description on the dataset page: https://huggingface.co/datasets/rakshi719/SpokenCOCO-A2IT.

sourceHugging Facecc-by-4.0updated 17d agoView on Hugging Face
0likes119downloads
settings

This repository belongs to rakshi719 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSpokenCOCO-A2IT
visibilitypublic
licencecc-by-4.0
gatedno
ownerrakshi719
Account settings
rakshi719/SpokenCOCO-A2IT · CoolFace