CoolFace
Datasetpublic

Wissam42/SEA-VL-Crawling-T2I

SeaVLCrawlingT2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are captions; the corpus contains images (text→image retrieval). Task category… See the full description on the dataset page: https://huggingface.co/datasets/Wissam42/SEA-VL-Crawling-T2I.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes34downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Wissam42/SEA-VL-Crawling-T2I · CoolFace