CoolFace
Datasetpublic

whybe-choi/SpokenCOCOA2IRetrieval

SpokenCOCOA2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SpokenCOCO pairs MS COCO images with recordings of human speakers reading the corresponding English captions. This task uses the 5,000-image Karpathy test split with 25,031 spoken captions. Queries are spoken captions and the corpus contains images; the goal is to retrieve the image described by each recording. Task category Any2AnyRetrieval (audio-to-image) Domains Scene, Spoken Reference… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/SpokenCOCOA2IRetrieval.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes90downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
whybe-choi/SpokenCOCOA2IRetrieval · CoolFace