artist/glami-1m-t2i-mteb
GLAMI-1M text-to-image retrieval This MTEB-formatted derivative uses the complete 116,004-row official GLAMI-1M test split. Product names and descriptions are text queries and product images are the corpus. Repeated image IDs and exact repeated texts are deduplicated within each language, and qrels retain every observed text-image association. The unchanged source archives are already hosted by the original authors in glami/glami-1m. GLAMI-1M-dataset--test-only.zip is pinned at… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-t2i-mteb.
GLAMI-1M text-to-image retrieval
This MTEB-formatted derivative uses the complete 116,004-row official GLAMI-1M test split. Product names and descriptions are text queries and product images are the corpus. Repeated image IDs and exact repeated texts are deduplicated within each language, and qrels retain every observed text-image association.
The unchanged source archives are already hosted by the original authors in `glami/glami-1m`. GLAMI-1M-dataset--test-only.zip is pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c with SHA-256 814c1fb456b86a1a4f1c44fe3fb15c6bc645480e70ea4bfe06bd6a76e29afe9b.
Each of the 13 language subsets has standard <lang>-queries, <lang>-corpus, and <lang>-qrels configurations with a test split.
