whybe-choi/SpokenCOCOA2IRetrieval
SpokenCOCOA2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SpokenCOCO pairs MS COCO images with recordings of human speakers reading the corresponding English captions. This task uses the 5,000-image Karpathy test split with 25,031 spoken captions. Queries are spoken captions and the corpus contains images; the goal is to retrieve the image described by each recording. Task category Any2AnyRetrieval (audio-to-image) Domains Scene, Spoken Reference… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/SpokenCOCOA2IRetrieval.
090
