CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vidore /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K1 likes2.1k downloads1y agoHugging Face02vidore /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes2k downloads1y agoHugging Face03vidore /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes1.9k downloads1y agoHugging Face04vidore /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.9k downloads1y agoHugging Face05mteb /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.2k downloads8mo agoHugging Face06mteb /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.1k downloads8mo agoHugging Face07mteb /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1k downloads8mo agoHugging Face08mteb /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes1k downloads8mo agoHugging Face09vidore /docvqa_test_subsampled Dataset Description This is the test set taken from the DocVQA dataset. It includes collected images from the UCSF Industry Documents Library. Questions and answers were manually annotated. Example of data (see viewer) Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs and renamed the different columns. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/docvqa_test_subsampled.imagedocument-question-answeringn<1K6 likes886 downloads1y agoHugging Face10vidore /tabfquad_test_subsampled Dataset Description TabFQuAD (Table French Question Answering Dataset) is designed to evaluate TableQA models in realistic industry settings. Using a vision language model (GPT4V), we create additional queries to augment the existing human-annotated ones. Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 280 pairs, leaving the rest for training and renaming the different columns. Load the dataset from… See the full description on the dataset page: https://huggingface.co/datasets/vidore/tabfquad_test_subsampled.imagedocument-question-answeringn<1K0 likes739 downloads1y agoHugging Face11vidore /arxivqa_test_subsampled Dataset Description This is a VQA dataset based on figures extracted from arXiv publications taken from ArXiVQA dataset from Multimodal ArXiV. The questions were generated synthetically using GPT-4 Vision. Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs. Furthermore we renamed the different columns for our purpose. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/arxivqa_test_subsampled.imagedocument-question-answeringn<1K4 likes648 downloads1y agoHugging Face12vidore /infovqa_test_subsampled Dataset Description This is the test set taken from the InfoVQA dataset. includes infographics collected from the Internet using the search query “infographics”. Questions and answers were manually annotated. Questions and answers were manually annotated. Example of data : (see viewer) Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs and renamed the different columns. Load the dataset from… See the full description on the dataset page: https://huggingface.co/datasets/vidore/infovqa_test_subsampled.imagedocument-question-answeringn<1K3 likes635 downloads1y agoHugging Face13justram /atomic_subsampled_500kimage100K<n<1M1 likes261 downloads2y agoHugging Face14vidore /docvqa_test_subsampled_tesseractimagedocument-question-answeringn<1K0 likes69 downloads1y agoHugging Face15vidore /tabfquad_test_subsampled_ocr_chunkimagen<1K0 likes56 downloads2y agoHugging Face16vidore /arxivqa_test_subsampled_tesseractimagedocument-question-answeringn<1K0 likes52 downloads1y agoHugging Face17andreaparker /wiki-screenshot-corpus_subsampled-5kimage1K<n<10K0 likes42 downloads2y agoHugging Face18vidore /infovqa_test_subsampled_tesseractimagedocument-question-answeringn<1K0 likes40 downloads1y agoHugging Face19vidore /tabfquad_test_subsampled_tesseractimagedocument-question-answeringn<1K0 likes38 downloads1y agoHugging Face20vidore /arxivqa_test_subsampled_captioningimagedocument-question-answering1K<n<10K1 likes36 downloads1y agoHugging Face21ManukyanD /MMEB-train-subsampledimage100K<n<1M0 likes28 downloads2y agoHugging Face22vidore /docvqa_test_subsampled_captioningimagedocument-question-answering1K<n<10K0 likes25 downloads1y agoHugging Face23vidore /docvqa_test_subsampled_ocr_chunkimage1K<n<10K0 likes23 downloads2y agoHugging Face24vidore /infovqa_test_subsampled_ocr_chunkimage1K<n<10K0 likes23 downloads2y agoHugging Face25vidore /arxivqa_test_subsampled_ocr_chunkimage1K<n<10K0 likes21 downloads2y agoHugging Face26nixiieee /AMAZON-Products-2023-subsampleimage1K<n<10K0 likes14 downloads11mo agoHugging Face27aakash-projects /arxivqa_test_subsampled Dataset Description This is a VQA dataset based on figures extracted from arXiv publications taken from ArXiVQA dataset from Multimodal ArXiV. The questions were generated synthetically using GPT-4 Vision. Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs. Furthermore we renamed the different columns for our purpose. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/arxivqa_test_subsampled.imagedocument-question-answeringn<1K0 likes14 downloads4mo agoHugging Face28vidore /tabfquad_test_subsampled_captioningimagedocument-question-answeringn<1K0 likes13 downloads1y agoHugging Face29vidore /infovqa_test_subsampled_captioningimagedocument-question-answering1K<n<10K0 likes13 downloads1y agoHugging Face30CATIE-AQ /retriever-vidore-tabfquad_test_subsampled-clean Description vidore/tabfquad_test_subsampled dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives. Citation @misc{faysse2024colpaliefficientdocumentretrieval, title={ColPali: Efficient Document Retrieval with Vision Language Models}, author={Manuel Faysse… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-vidore-tabfquad_test_subsampled-clean.imageimage-feature-extractionn<1K0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.