datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Computer-Science-Pretrainingreproduction_qwen235b_computersciencereproduction_o4mini_computerscienceyourbench_reproduction_o4mini_computersciencereproduction_deepseekr1_computersciencereproduction_g3_mini_computerscienceViet-ComputerScience-VQA
Dataset Overview
This dataset is was created from 6899 Vietnamese 🇻🇳 Computer Science books. Each image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 40,000 detailed descriptions, and query-based questions and answers generated by the Gemini 1.5 Flash model, currently Google's leading model on the WildVision Arena Leaderboard. This results in a richly annotated dataset, ideal for… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-ComputerScience-VQA.original_mmlu_pro_computerscience
