datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GeneralScience-MLLM-22K
GeneralScience-MLLM-22K
Dataset Summary
GeneralScience-MLLM-22K is a unified general-science multiple-choice QA collection built from local snapshots of SciQ, AI2 ARC, and ScienceQA. It follows the same release style as a subject-specific MLLM dataset: every sample is stored as one JSONL record, text-only and image-text examples share one schema, and ScienceQA images are exported as standalone files referenced by relative paths.
The release contains 22,661… See the full description on the dataset page: https://huggingface.co/datasets/gineven/GeneralScience-MLLM-22K.General-Evol-VQA
Dataset Card for General-Evol-VQA-1.2M
This dataset has been carefully curated to enhance the general instruction capabilities of Vision-Language Models (VLMs). It comprises two subsets:
600k English samples
600k Korean samples
We recommend using this dataset alongside other task-specific datasets (e.g., OCR, Language, code, math, ...) to improve performance and achieve more robust model capabilities.
Made by: maum.ai Brain NLP. Jaeyoon Jung, Yoonshik Kim
Dataset Target… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/General-Evol-VQA.GeneralHistoryOfAfricaXI
GeneralHistoryOfAfricaXI
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
3731
Avg chars/chunk
696
Avg images/chunk
0.00
Source files
2
Duplicates removed
1
Quality filtered
70
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned text… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/GeneralHistoryOfAfricaXI.
