CoolFace
Datasetpublic

whybe-choi/SpokenCOCOA2IRetrieval

SpokenCOCOA2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SpokenCOCO pairs MS COCO images with recordings of human speakers reading the corresponding English captions. This task uses the 5,000-image Karpathy test split with 25,031 spoken captions. Queries are spoken captions and the corpus contains images; the goal is to retrieve the image described by each recording. Task category Any2AnyRetrieval (audio-to-image) Domains Scene, Spoken Reference… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/SpokenCOCOA2IRetrieval.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes90downloads
10 commits on main
72731c42mo ago

Add default-qrels

whybe-choi
c55f3c42mo ago

Add corpus-corpus

whybe-choi
4060e082mo ago

Add queries-queries

whybe-choi
b74eb902mo ago

Add descriptive statistics

whybe-choi
0b26c682mo ago

Add dataset card

whybe-choi
8a9727f2mo ago

Add dataset card

whybe-choi
f9e5e4b2mo ago

Add default-qrels

whybe-choi
f8fe0d62mo ago

Add corpus-corpus

whybe-choi
daf73892mo ago

Add queries-queries

whybe-choi
95c4f062mo ago

initial commit

whybe-choi