datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ammluAMMLU is a comprehensive Arabic assessment suite specifically designed to evaluate the advanced knowledge and reasoning abilities of LLMs within the Arabic language and cultural context.A-MMK12-8K
A-MMK12-8K
This dataset is created using the SynthRL pipeline to synthesize 3,380 challenging questions from 8,072 seed samples.
Dataset Details
Synthesis Method: SynthRL pipeline
Total Samples: 11,452 (8,072 seed + 3,380 synthesized)
Seed Data: MMK12 dataset
Purpose: Training data for VLM reinforcement learning with verifiable rewards (RLVR)
Data Sources
Original MMK12: Proposed in MM-EUREKA (thanks to the MM-EUREKA authors)
Processed Seed Data: Our… See the full description on the dataset page: https://huggingface.co/datasets/Jakumetsu/A-MMK12-8K.Maths-Grade-SchoolMaths-Grade-School
I am releasing large Grade School level Mathematics datatset.
This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a diverse array of topics fundamental to building a strong mathematical foundation.
This dataset is in instruction format so that model developers, researchers etc. can easily use this dataset.
Following Fields & sub Fields are covered:
Calculus
Probability
Algebra
Liner Algebra
Trigonometry
Differential Equations… See the full description on the dataset page: https://huggingface.co/datasets/ammulu03/Maths-Grade-School.
