datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Finch-Collection-Gemini-3-Flash
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions.
👋 This is the Gemini-3-Flash teacher variant of the Finch Collection — evolutionary search trajectories from the paper Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks, but with Gemini-3-Flash as the teacher mutation… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection-Gemini-3-Flash.FLAWS
FLAWS: Faults Localization Across Writing in Science
FLAWS is a benchmark for evaluating error identification and localization in scientific papers. It currently consists of 713 paper–error examples, including:
265 unique papers with one error inserted using GPT-5 (in ALL_OPENAI.tar.gz)
448 unique papers with one error inserted using Gemini 2.5 Pro (in ALL_GEMINI.tar.gz)
The dataset is generated using a systematic, autonomous framework that produces paper–error examples and… See the full description on the dataset page: https://huggingface.co/datasets/xasayi/FLAWS.bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Some random images with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Viet-Doc-VQA-II-flash2
Dataset Overview
This dataset is a continuation of the ongoing work from Viet Document VAQ dataset was collected from 64,765 pages of Vietnamese 🇻🇳 textbooks( Sách bài tập, chuyên đề, sách giáo án của Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 388,277 detailed… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-II-flash2.Viet-OCR-VQA-flash2
Dataset Overview
The dataset comprises over 137,000 images potentially containing Vietnamese 🇻🇳 textual content. It was curated using the Gemini 1.5 Flash model, currently Google model leading on the WildVision Arena Leaderboard for Visual Question Answering (VQA). Each image is accompanied by a detailed description and 5 self-generated questions and answers related to the textual content within the image.
In total, there are more than 822,679 individual questions, encompassing… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-OCR-VQA-flash2.Viet-Doc-VQA-flash2
Dataset Overview
The Document VAQ dataset was collected from 51,856 pages of Vietnamese 🇻🇳 textbooks( Sách Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 310,952 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-flash2.
