datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
k12-freeformMM-K12
MM-K12
[📂 GitHub] [📜 Paper]
MM-K12 is a curated, high-quality dataset containing 10,000 multimodal math problems sourced from K-12 educational content. Each problem includes both textual and visual components, covering a wide range of mathematical topics (e.g., arithmetic, geometry, algebra). All problems have unique, verifiable answers, making the dataset ideal for supervised training, evaluation, and reward modeling in multimodal mathematical reasoning tasks.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Cierra0506/MM-K12.k12-freeform-extendedk12K12The train set is K12. The test set is MathVista.
k12_printing_cleaned
k12_printing_cleaned
The k12_printing__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
214,293
QA turns
501,016
answers rewritten by the cleaning pass
104,052
QA created by the cleaning pass (new_qa)
286,819 (57.2%)
shards
8
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/k12_printing_cleaned.K12_2k1filtered_k12_resample_no_chineseK12-QwenMM-K12K12-Freeformk12-resamplefiltered_k12_resample_no_chinese
