CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pixparse /docvqa-single-page-questions Dataset Card for DocVQA Dataset Dataset Summary DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images. Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information. Usage This dataset can be used with current releases of Hugging Face datasets library. Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.imagequestion-answering10K<n<100K11 likes2.8k downloads2y agoHugging Face02dfkiuser /kangaroo_math_mc_questionsimage1K<n<10K0 likes1.1k downloads8mo agoHugging Face03Jiiwonn /roco2-question-id-dataset ROCOv2: Radiology Object in COntext version 2 Introduction ROCOv2 is a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PMC Open Access Subset. It is an updated version of the ROCO dataset, adding 35,705 new images and improving concept extraction and filtering. Dataset Overview The ROCOv2 dataset contains 79,789 radiological images, each with a corresponding caption and medical concepts. The… See the full description on the dataset page: https://huggingface.co/datasets/Jiiwonn/roco2-question-id-dataset.image10K<n<100K0 likes305 downloads1y agoHugging Face04alakxender /dhivehi-vrd-batch-1-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes216 downloads1y agoHugging Face05OneEyeDJ /Art-Vision-Question-Answering-Dataset Art Vision Question Answering Dataset 🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering. Dataset Overview This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks. ✨ Key Features 🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer 💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.imageimage-to-textn<1K2 likes187 downloads1y agoHugging Face06alakxender /dhivehi-vrd-batch-3-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes137 downloads1y agoHugging Face07alakxender /dhivehi-vrd-batch-6-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes135 downloads1y agoHugging Face08taesiri /video-game-question-answeringimage10K<n<100K3 likes134 downloads3y agoHugging Face09Jiiwonn /rocov2-questions-radiologyimage10K<n<100K1 likes113 downloads1y agoHugging Face10chainyo /rvl-cdip-questionnaire⚠️ This only a subpart of the original dataset, containing only questionnaire. The RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class. There are 320,000 training images, 40,000 validation images, and 40,000 test images. The images are sized so their largest dimension does not exceed 1000 pixels. For questions and comments please contact Adam Harley (aharley@scs.ryerson.ca). The full… See the full description on the dataset page: https://huggingface.co/datasets/chainyo/rvl-cdip-questionnaire.image10K<n<100K0 likes104 downloads4y agoHugging Face11alakxender /dhivehi-vrd-batch-2-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes100 downloads1y agoHugging Face12fawern /visual-question-answering-cocoimagen<1K6 likes98 downloads2y agoHugging Face13moteloumka /movie-frames-questionsimage1K<n<10K0 likes97 downloads2y agoHugging Face14zipu-w /GaoKao-questionsimagen<1K0 likes87 downloads1y agoHugging Face15alakxender /dhivehi-vrd-batch-4-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes85 downloads1y agoHugging Face16alakxender /dhivehi-vrd-batch-5-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes78 downloads1y agoHugging Face17indrehus /docvqa-single-page-questions-answer-ocrgated DocVQA with Answer Localization This dataset provides answer-localization annotations produced by our pipeline on top of the DocVQA dataset. Usage from datasets import load_dataset # Load the dataset with answer OCR annotations ds = load_dataset("indrehus/docvqa-single-page-questions-answer-ocr", split="validation") # Get a single sample sample = ds[0] # Available fields in each sample: print("Image:", sample["image"]) # PIL.Image print("Question:"… See the full description on the dataset page: https://huggingface.co/datasets/indrehus/docvqa-single-page-questions-answer-ocr.imagequestion-answering10K<n<100K0 likes77 downloads5mo agoHugging Face18Berkesule /translated_visual_puzzles_with_questionimagen<1K0 likes69 downloads9mo agoHugging Face19mashey /dhivehi-vrd-batch-2-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text100K<n<1M0 likes44 downloads4mo agoHugging Face20tejasexpress /question-answerimagen<1K1 likes42 downloads3y agoHugging Face21mikeogezi /vsr_random_with_questionsimage10K<n<100K0 likes36 downloads2y agoHugging Face22Jiiwonn /roco2-question-dataset-validationimage1K<n<10K0 likes31 downloads1y agoHugging Face23geekyrakshit /indian-exam-questionsimage1K<n<10K0 likes30 downloads3mo agoHugging Face24mikeogezi /vsr_zeroshot_with_questionsimage1K<n<10K0 likes28 downloads2y agoHugging Face25alakxender /dhivehi-vrd-batch-7-img-questions Dhivehi Single-Line Text-Image Dataset A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row). Note: This dataset is a subset from alakxender/dhivehi-vrd-images. imageimage-to-text10K<n<100K0 likes28 downloads1y agoHugging Face26zipu-w /alevel-exam-questionsimagen<1K0 likes22 downloads1y agoHugging Face27munteanr20 /driving-benchmark-questionsgatedimage1K<n<10K0 likes19 downloads3mo agoHugging Face28zipu-w /AP-exam-questionsimagen<1K0 likes18 downloads1y agoHugging Face29Berkesule /translated_mmiq_dataset_with_questionimage1K<n<10K0 likes18 downloads10mo agoHugging Face30raynardj /cauldron-questionimage1K<n<10K0 likes18 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.