CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /flickr30ki2timage10K<n<100K0 likes6.9k downloads2y agoHugging Face02mteb /xm3600 XM3600T2IRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve images based on multilingual descriptions. Task category Any2AnyMultilingualRetrieval (text-to-image) Domains Encyclopaedic, Written Reference Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Source datasets: mteb/xm3600 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import… See the full description on the dataset page: https://huggingface.co/datasets/mteb/xm3600.imagevisual-document-retrieval100K<n<1M0 likes1.8k downloads4mo agoHugging Face03artist /mvl-sib-sent2img-mteb MVL-SIB sentence-to-image for MTEB Native MTEB retrieval packaging of the official MVL-SIB single-reference (k=1) sentence-to-image task. Each of 205 language subsets has 3,012 sentence queries, four candidate images per query, and one correct image. All subsets reference one shared 70-image corpus file. Source and changes Derived from the official WueNLP/MVL-SIB dataset and MVL-SIB paper, pinned at 1df5974e8fb204e91ee70cef2b3b7196a14b390f. The official builder's… See the full description on the dataset page: https://huggingface.co/datasets/artist/mvl-sib-sent2img-mteb.image1M<n<10M0 likes1.2k downloads24d agoHugging Face04mteb /cifar10 Dataset Card for CIFAR-10 Dataset Summary The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cifar10.imageimage-classification10K<n<100K3 likes1.2k downloads8mo agoHugging Face05mteb /tatdqa_test_beirBEIR version of vidore/tatdqa_test. imagedocument-question-answering1K<n<10K0 likes1.1k downloads7mo agoHugging Face06mteb /imagenet-dog-15image1K<n<10K0 likes1.1k downloads8mo agoHugging Face07mteb /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.1k downloads7mo agoHugging Face08vidore /vidore_v3_energy_mteb_format Vidore3EnergyRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_energy How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3EnergyRetrieval") evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_energy_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes1.1k downloads11mo agoHugging Face09mteb /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.1k downloads7mo agoHugging Face10vidore /vidore_v3_computer_science_mteb_format Vidore3ComputerScienceRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_computer_science How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes1.1k downloads11mo agoHugging Face11mteb /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes1k downloads7mo agoHugging Face12vidore /vidore_v3_finance_en_mteb_format Vidore3FinanceEnRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_en How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceEnRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_en_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes1k downloads11mo agoHugging Face13mteb /sun397image100K<n<1M0 likes995 downloads8mo agoHugging Face14mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes995 downloads7mo agoHugging Face15mteb /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes980 downloads7mo agoHugging Face16mteb /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes967 downloads7mo agoHugging Face17vidore /vidore_v3_hr_mteb_format Vidore3HrRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_hr How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3HrRetrieval") evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_hr_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes952 downloads11mo agoHugging Face18mteb /shiftproject_test_beirBEIR version of vidore/shiftproject_test. imagedocument-question-answering1K<n<10K0 likes947 downloads7mo agoHugging Face19mteb /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K0 likes947 downloads7mo agoHugging Face20mteb /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes944 downloads7mo agoHugging Face21vidore /vidore_v3_industrial_mteb_format Vidore3IndustrialRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_industrial How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3IndustrialRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes936 downloads11mo agoHugging Face22vidore /vidore_v3_pharmaceuticals_mteb_format Vidore3PharmaceuticalsRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_pharmaceuticals How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes890 downloads11mo agoHugging Face23vidore /vidore_v3_finance_fr_mteb_format Vidore3FinanceFrRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_fr How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceFrRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_fr_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes881 downloads11mo agoHugging Face24vidore /vidore_v3_physics_mteb_format Vidore3PhysicsRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_physics How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3PhysicsRetrieval") evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_physics_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes832 downloads11mo agoHugging Face25mteb /forb_retrievalimage10K<n<100K0 likes830 downloads2y agoHugging Face26mteb /gld-v2image100K<n<1M0 likes794 downloads7mo agoHugging Face27mteb /JinaVDRWikimediaCommonsDocumentsRetrievalimage10K<n<100K0 likes656 downloads11mo agoHugging Face28jupyterjazz /XModBench-MTEB XModBench-Lite for MTEB This repository is a deterministic MTEB normalization of the official RyanWW/XModBench XModBench-Lite release at revision a679188cf062b9810d2e09c2edabc0b1aef9f244. The source contains 6,000 four-choice questions balanced across six canonical modality configurations and five capability families. This MTEB adaptation retains 5,981 questions. It excludes 19 questions that reference five unusable MP4 files in the pinned official archive. Four are truncated:… See the full description on the dataset page: https://huggingface.co/datasets/jupyterjazz/XModBench-MTEB.audio10K<n<100K0 likes564 downloads25d agoHugging Face29mteb /XM3600T2IRetrieval XM3600T2IRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve images based on multilingual descriptions. Task category t2i Domains Encyclopaedic, Written Reference https://aclanthology.org/2022.emnlp-main.45/ Source datasets: floschne/xm3600 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("XM3600T2IRetrieval") evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/XM3600T2IRetrieval.imagevisual-document-retrieval100K<n<1M0 likes537 downloads11mo agoHugging Face30mteb /Vidore3IndustrialOCRRetrieval Vidore3IndustrialOCRRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. This dataset, Industrial reports, is a corpus of technical documents on military aircraft (fueling, mechanics...), intended for complex-document understanding tasks. Original queries were created in english, then translated to french, german, italian, portuguese and spanish. This variant includes the OCR'ed markdown so allow for comparison across image-text… See the full description on the dataset page: https://huggingface.co/datasets/mteb/Vidore3IndustrialOCRRetrieval.imagevisual-document-retrieval10K<n<100K0 likes467 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.