datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hero_run_4_math_code32b_exploit_seed_math_code_dedup_decontaminateJosephgflowers__TinyLlama_v1.1_math_code-world-test-1-details
Dataset Card for Evaluation run of Josephgflowers/TinyLlama_v1.1_math_code-world-test-1
Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama_v1.1_math_code-world-test-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details.math-ai-bench-sources-code
math-ai-bench-sources-code
A code benchmark evaluation dataset with 83,072 solution trajectories generated by state-of-the-art thinking models on coding benchmark problems.
Overview
Each entry is a long-form solution trajectory (chain-of-thought + final code) produced by a reasoning model on a held-out coding benchmark. Every trajectory carries a verified correct label, and every problem carries a correct_ratio (pass rate over all trajectories for that problem).… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources-code.DeepMount00__Qwen2.5-7B-Instruct-MathCoder-details
Dataset Card for Evaluation run of DeepMount00/Qwen2.5-7B-Instruct-MathCoder
Dataset automatically created during the evaluation run of model DeepMount00/Qwen2.5-7B-Instruct-MathCoder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepMount00__Qwen2.5-7B-Instruct-MathCoder-details.EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.Qwen2.5-7B-SFT-Math-Code-1M-1000-MATH500mmlu-olmo37b-stage2-math-code-analysis-with-contextmmlu-olmo37b-stage2-math-code-analysis-baselineSYNTHETIC-2-SFT-cn-fltrd-final-ngram-filtered-chinese-filtered-math-only-no-codeQwen2.5-7B-SFT-Math-Code-1M-AIME25dolma_20bn_no_math_codeQwen__Qwen2.5-Coder-32B-InstructG-TasksEpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPOg-tasks-2g-tasks-3gsm8k___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructproofwriter___3txt___Qwen2.5_7B_Instruct___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_InstructQwen3-30B-A3B-SFT-Math-Code-1M-1000-MATH500proofwriter___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_InstructQwen__Qwen2.5-Coder-14B-InstructQwen__Qwen2.5-Coder-7B-InstructEval-Tasksmath500___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructmath500___3txt___Qwen2.5_7B_Instruct___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instruct
