CoolFace
Datasetpublic

whybe-choi/SpokenCOCOA2IRetrieval

SpokenCOCOA2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SpokenCOCO pairs MS COCO images with recordings of human speakers reading the corresponding English captions. This task uses the 5,000-image Karpathy test split with 25,031 spoken captions. Queries are spoken captions and the corpus contains images; the goal is to retrieve the image described by each recording. Task category Any2AnyRetrieval (audio-to-image) Domains Scene, Spoken Reference… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/SpokenCOCOA2IRetrieval.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes90downloads

whybe-choi/SpokenCOCOA2IRetrieval · main · files are served by the source, never re-hosted here