CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cmarkea /table-vqa Dataset description The table-vqa Dataset integrates images of tables from the dataset AFTdb (Arxiv Figure Table Database) curated by cmarkea. This dataset consists of pairs of table images and corresponding LaTeX source code, with each image linked to an average of ten questions and answers. Half of the Q&A pairs are in English and the other half in French. These questions and answers were generated using Gemini 1.5 Pro and Claude 3.5 sonnet, making the dataset well-suited for… See the full description on the dataset page: https://huggingface.co/datasets/cmarkea/table-vqa.imagetext-generation10K<n<100K24 likes501 downloads2y agoHugging Face02LR-AI-Labs /vi-OCR_VQA Dataset Card for "vi-OCR-VQA" imagevisual-question-answering10K<n<100K7 likes61 downloads2y agoHugging Face035CD-AI /Viet-Doc-VQA-II-flash2gated Dataset Overview This dataset is a continuation of the ongoing work from Viet Document VAQ dataset was collected from 64,765 pages of Vietnamese 🇻🇳 textbooks( Sách bài tập, chuyên đề, sách giáo án của Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset. There is a set of 388,277 detailed… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-II-flash2.imagevisual-question-answering10K<n<100K6 likes21 downloads8mo agoHugging Face045CD-AI /Viet-OCR-VQA-flash2gated Dataset Overview The dataset comprises over 137,000 images potentially containing Vietnamese 🇻🇳 textual content. It was curated using the Gemini 1.5 Flash model, currently Google model leading on the WildVision Arena Leaderboard for Visual Question Answering (VQA). Each image is accompanied by a detailed description and 5 self-generated questions and answers related to the textual content within the image. In total, there are more than 822,679 individual questions, encompassing… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-OCR-VQA-flash2.imagevisual-question-answering100K<n<1M8 likes17 downloads8mo agoHugging Face05AhmadIshaqai /brain-vqa-radimagequestion-answeringn<1K0 likes15 downloads1y agoHugging Face065CD-AI /Viet-Doc-VQA-flash2gated Dataset Overview The Document VAQ dataset was collected from 51,856 pages of Vietnamese 🇻🇳 textbooks( Sách Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset. There is a set of 310,952 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-flash2.imagevisual-question-answering10K<n<100K4 likes13 downloads8mo agoHugging Face07odiagenmllm /odia_vqa_en_odi_setgated Dataset Card for OVQA Instruction Set Dataset Summary Odia Visual Question Answering (OVQA) Instruction Set is a multimodal dataset comprising text and images structured in an instruction format, designed for developing Multimodal Large Language Models (MLLMs). Supported Tasks and Leaderboards Multimodal Large Language Model (MLLM) Languages Odia, English Dataset Structure JSON Paper For more details on data preparation… See the full description on the dataset page: https://huggingface.co/datasets/odiagenmllm/odia_vqa_en_odi_set.imagetext-generation10K<n<100K3 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.