CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /flickr30ki2timage10K<n<100K0 likes7k downloads2y agoHugging Face02mteb /xm3600 XM3600T2IRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve images based on multilingual descriptions. Task category Any2AnyMultilingualRetrieval (text-to-image) Domains Encyclopaedic, Written Reference Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Source datasets: mteb/xm3600 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import… See the full description on the dataset page: https://huggingface.co/datasets/mteb/xm3600.imagevisual-document-retrieval100K<n<1M0 likes1.8k downloads4mo agoHugging Face03artist /mvl-sib-sent2img-mteb MVL-SIB sentence-to-image for MTEB Native MTEB retrieval packaging of the official MVL-SIB single-reference (k=1) sentence-to-image task. Each of 205 language subsets has 3,012 sentence queries, four candidate images per query, and one correct image. All subsets reference one shared 70-image corpus file. Source and changes Derived from the official WueNLP/MVL-SIB dataset and MVL-SIB paper, pinned at 1df5974e8fb204e91ee70cef2b3b7196a14b390f. The official builder's… See the full description on the dataset page: https://huggingface.co/datasets/artist/mvl-sib-sent2img-mteb.image1M<n<10M0 likes1.2k downloads24d agoHugging Face04mteb /cifar10 Dataset Card for CIFAR-10 Dataset Summary The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cifar10.imageimage-classification10K<n<100K3 likes1.2k downloads8mo agoHugging Face05mteb /tatdqa_test_beirBEIR version of vidore/tatdqa_test. imagedocument-question-answering1K<n<10K0 likes1.1k downloads7mo agoHugging Face06mteb /imagenet-dog-15image1K<n<10K0 likes1.1k downloads8mo agoHugging Face07mteb /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.1k downloads7mo agoHugging Face08mteb /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.1k downloads7mo agoHugging Face09mteb /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes1k downloads7mo agoHugging Face10mteb /sun397image100K<n<1M0 likes996 downloads8mo agoHugging Face11vidore /vidore_v3_computer_science_mteb_format Vidore3ComputerScienceRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_computer_science How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes995 downloads11mo agoHugging Face12mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes992 downloads7mo agoHugging Face13vidore /vidore_v3_finance_en_mteb_format Vidore3FinanceEnRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_en How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceEnRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_en_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes990 downloads11mo agoHugging Face14mteb /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes981 downloads7mo agoHugging Face15mteb /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes958 downloads7mo agoHugging Face16mteb /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K0 likes948 downloads7mo agoHugging Face17mteb /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes946 downloads7mo agoHugging Face18mteb /shiftproject_test_beirBEIR version of vidore/shiftproject_test. imagedocument-question-answering1K<n<10K0 likes945 downloads7mo agoHugging Face19vidore /vidore_v3_industrial_mteb_format Vidore3IndustrialRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_industrial How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3IndustrialRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes911 downloads11mo agoHugging Face20vidore /vidore_v3_hr_mteb_format Vidore3HrRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_hr How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3HrRetrieval") evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_hr_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes891 downloads11mo agoHugging Face21vidore /vidore_v3_pharmaceuticals_mteb_format Vidore3PharmaceuticalsRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_pharmaceuticals How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes838 downloads11mo agoHugging Face22vidore /vidore_v3_finance_fr_mteb_format Vidore3FinanceFrRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_fr How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceFrRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_fr_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes836 downloads11mo agoHugging Face23mteb /forb_retrievalimage10K<n<100K0 likes831 downloads2y agoHugging Face24vidore /vidore_v3_physics_mteb_format Vidore3PhysicsRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_physics How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3PhysicsRetrieval") evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_physics_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes805 downloads11mo agoHugging Face25vidore /vidore_v3_energy_mteb_format Vidore3EnergyRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_energy How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3EnergyRetrieval") evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_energy_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes797 downloads11mo agoHugging Face26mteb /gld-v2image100K<n<1M0 likes741 downloads7mo agoHugging Face27mteb /JinaVDRWikimediaCommonsDocumentsRetrievalimage10K<n<100K0 likes656 downloads11mo agoHugging Face28jupyterjazz /XModBench-MTEB XModBench-Lite for MTEB This repository is a deterministic MTEB normalization of the official RyanWW/XModBench XModBench-Lite release at revision a679188cf062b9810d2e09c2edabc0b1aef9f244. The source contains 6,000 four-choice questions balanced across six canonical modality configurations and five capability families. This MTEB adaptation retains 5,981 questions. It excludes 19 questions that reference five unusable MP4 files in the pinned official archive. Four are truncated:… See the full description on the dataset page: https://huggingface.co/datasets/jupyterjazz/XModBench-MTEB.audio10K<n<100K0 likes534 downloads25d agoHugging Face29mteb /XM3600T2IRetrieval XM3600T2IRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve images based on multilingual descriptions. Task category t2i Domains Encyclopaedic, Written Reference https://aclanthology.org/2022.emnlp-main.45/ Source datasets: floschne/xm3600 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("XM3600T2IRetrieval") evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/XM3600T2IRetrieval.imagevisual-document-retrieval100K<n<1M0 likes533 downloads11mo agoHugging Face30mteb /Vidore3IndustrialOCRRetrieval Vidore3IndustrialOCRRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. This dataset, Industrial reports, is a corpus of technical documents on military aircraft (fueling, mechanics...), intended for complex-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish. This variant includes the OCR'ed markdown so allow for comparison across image-text… See the full description on the dataset page: https://huggingface.co/datasets/mteb/Vidore3IndustrialOCRRetrieval.imagevisual-document-retrieval10K<n<100K0 likes478 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.