CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations-dev /hero_run_4_math_codetabular1M<n<10M0 likes374 downloads1y agoHugging Face02mlfoundations-dev /32b_exploit_seed_math_code_dedup_decontaminatetabular100K<n<1M0 likes170 downloads2y agoHugging Face03open-llm-leaderboard /Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-detailsgated Dataset Card for Evaluation run of Josephgflowers/TinyLlama_v1.1_math_code-world-test-1 Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama_v1.1_math_code-world-test-1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details.tabular10K<n<100K0 likes61 downloads2y agoHugging Face04haowu89 /math-ai-bench-sources-code math-ai-bench-sources-code A code benchmark evaluation dataset with 83,072 solution trajectories generated by state-of-the-art thinking models on coding benchmark problems. Overview Each entry is a long-form solution trajectory (chain-of-thought + final code) produced by a reasoning model on a held-out coding benchmark. Every trajectory carries a verified correct label, and every problem carries a correct_ratio (pass rate over all trajectories for that problem).… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources-code.tabular10K<n<100K0 likes45 downloads3mo agoHugging Face05open-llm-leaderboard /DeepMount00__Qwen2.5-7B-Instruct-MathCoder-detailsgated Dataset Card for Evaluation run of DeepMount00/Qwen2.5-7B-Instruct-MathCoder Dataset automatically created during the evaluation run of model DeepMount00/Qwen2.5-7B-Instruct-MathCoder The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepMount00__Qwen2.5-7B-Instruct-MathCoder-details.tabular10K<n<100K0 likes39 downloads2y agoHugging Face06open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face07mikasenghaas /Qwen2.5-7B-SFT-Math-Code-1M-1000-MATH500tabularn<1K0 likes18 downloads1y agoHugging Face08SALT-NLP /mmlu-olmo37b-stage2-math-code-analysis-with-contexttabular1K<n<10K0 likes18 downloads2mo agoHugging Face09SALT-NLP /mmlu-olmo37b-stage2-math-code-analysis-baselinetabularn<1K0 likes16 downloads2mo agoHugging Face10saurabh5 /SYNTHETIC-2-SFT-cn-fltrd-final-ngram-filtered-chinese-filtered-math-only-no-codetabular10K<n<100K0 likes15 downloads1y agoHugging Face11mikasenghaas /Qwen2.5-7B-SFT-Math-Code-1M-AIME25tabular1K<n<10K0 likes15 downloads1y agoHugging Face12orionweller /dolma_20bn_no_math_codetabular10M<n<100M0 likes13 downloads2y agoHugging Face13math-extraction-comp /Qwen__Qwen2.5-Coder-32B-Instructtabular1K<n<10K0 likes13 downloads2y agoHugging Face14MathCodeBench /G-Taskstabularn<1K0 likes11 downloads2y agoHugging Face15math-extraction-comp /EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPOtabular1K<n<10K0 likes11 downloads2y agoHugging Face16MathCodeBench /g-tasks-2tabular1K<n<10K0 likes10 downloads2y agoHugging Face17MathCodeBench /g-tasks-3tabularn<1K0 likes10 downloads2y agoHugging Face18smoorsmith /gsm8k___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructtabular1K<n<10K0 likes10 downloads1y agoHugging Face19smoorsmith /proofwriter___3txt___Qwen2.5_7B_Instruct___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructtabularn<1K0 likes10 downloads1y agoHugging Face20mikasenghaas /Qwen3-30B-A3B-SFT-Math-Code-1M-1000-MATH500tabularn<1K0 likes10 downloads1y agoHugging Face21smoorsmith /proofwriter___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructtabularn<1K0 likes9 downloads1y agoHugging Face22math-extraction-comp /Qwen__Qwen2.5-Coder-14B-Instructtabular1K<n<10K0 likes6 downloads2y agoHugging Face23math-extraction-comp /Qwen__Qwen2.5-Coder-7B-Instructtabular1K<n<10K0 likes6 downloads2y agoHugging Face24MathCodeBench /Eval-Taskstabularn<1K0 likes4 downloads2y agoHugging Face25smoorsmith /math500___2txt___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructtabularn<1K0 likes4 downloads1y agoHugging Face26smoorsmith /math500___3txt___Qwen2.5_7B_Instruct___Qwen2.5_Math_7B_Instruct___Qwen2.5_Coder_7B_Instructtabularn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.