rakshi719/SpokenCOCO-A2IT
SpokenCOCO Audio-to-(Image+Text) Retrieval MTEB/MOEB task where queries are spoken audio captions and corpus items contain both a MSCOCO image and its written text caption. Task Given a spoken audio description of an image, retrieve the correct (image, text) pair from the corpus. Only models that can process all three modalities — audio, image, and text — can exploit the full corpus signal. Contents Queries: 25031 spoken audio captions (WAV… See the full description on the dataset page: https://huggingface.co/datasets/rakshi719/SpokenCOCO-A2IT.
This repository belongs to rakshi719 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
