datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hero_run_4_math_code32b_exploit_seed_math_code_dedup_decontaminatemath-ai-bench-sources-code
math-ai-bench-sources-code
A code benchmark evaluation dataset with 83,072 solution trajectories generated by state-of-the-art thinking models on coding benchmark problems.
Overview
Each entry is a long-form solution trajectory (chain-of-thought + final code) produced by a reasoning model on a held-out coding benchmark. Every trajectory carries a verified correct label, and every problem carries a correct_ratio (pass rate over all trajectories for that problem).… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources-code.Qwen2.5-7B-SFT-Math-Code-1M-1000-MATH500mmlu-olmo37b-stage2-math-code-analysis-with-contextmmlu-olmo37b-stage2-math-code-analysis-baselineSYNTHETIC-2-SFT-cn-fltrd-final-ngram-filtered-chinese-filtered-math-only-no-codeQwen2.5-7B-SFT-Math-Code-1M-AIME25dolma_20bn_no_math_codeQwen__Qwen2.5-Coder-32B-InstructG-TasksEpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPOg-tasks-2g-tasks-3gsm8k___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructproofwriter___3txt___Qwen2.5_7B_Instruct___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_InstructQwen3-30B-A3B-SFT-Math-Code-1M-1000-MATH500proofwriter___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_InstructQwen__Qwen2.5-Coder-14B-InstructQwen__Qwen2.5-Coder-7B-InstructEval-Tasksmath500___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructmath500___3txt___Qwen2.5_7B_Instruct___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instruct
