CoolFace
Datasetpublic

mteb/SEA-VL-Crawling-I2T

SeaVLCrawlingI2TRetrieval An MTEB dataset Massive Text Embedding Benchmark SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are images; the corpus contains captions (image→text retrieval). Task category… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SEA-VL-Crawling-I2T.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes34downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
mteb/SEA-VL-Crawling-I2T · CoolFace