datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math500math500math500-cot-deepseek-r1-1.5b
MATH-500 CoT completions (DeepSeek-R1-Distill-Qwen-1.5B)
Successful chain-of-thought completions for HuggingFaceH4/MATH-500 test problems, generated with deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B via vLLM.
Files
File
Description
records.parquet
Main dataset: correct completions as token IDs
manifest.json
Schema, tokenizer, run ids, decoding config
problem_index.json
unique_id → problem_idx in MATH-500 test
subject_max_tokens.json
Per-subject completion… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/math500-cot-deepseek-r1-1.5b.R-HORIZON-Math500R-HORIZON-Math500
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Math500.Qwen2.5-1.5B-Instruct-uPRM-70B-T80-math500-best_of_n-completionsLlama-3.2-1B-Instruct-uPRM-32B-T80-math500-best_of_n-completionsLlama-3.2-1B-Instruct-uPRM-70B-T80-math500-best_of_n-completionsQwen2.5-1.5B-Instruct-uPRM-32B-T80-math500-best_of_n-completionsm-math500mixed_sft_math500_128_s1_tulu2_sft_s1_1.0pctCustom-Bespoke-Stratos-7B_1753961306_eval_466d_math500_skip_attn_1
chengfu0118/Custom-Bespoke-Stratos-7B_1753961306_eval_466d_math500_skip_attn_1
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 1.80%
Accuracy
Questions Solved
Total Questions
1.80%
9
500
math500-bon-weighted-results
MATH-500 Best-of-N Weighted Selection Results
Dataset Description
This dataset contains the results of evaluating Best-of-N weighted selection on a subset of the MATH-500 benchmark. It was created as part of a HuggingFace internship exercise exploring how test-time compute scaling with reward models can improve LLM performance on math problems.
How It Was Constructed
1. Problem Selection
Started from the HuggingFaceH4/MATH-500 dataset (500 problems)… See the full description on the dataset page: https://huggingface.co/datasets/cmpatino/math500-bon-weighted-results.math500-typosmath500-checkMATH-500-gsm8k-format
MATH-500-gsm8k-format
Dataset Description
This dataset contains 500 mathematical problems from the MATH-500 benchmark, converted to GSM8K format for step-by-step reasoning.
Dataset Summary
Source: HuggingFaceH4/MATH-500
Format: GSM8K-style step-by-step solutions with inline computation annotations
Size: 500 problems
Split: Test (original MATH-500 test split)
Conversion Process
The original MATH-500 solutions (which use LaTeX notation and… See the full description on the dataset page: https://huggingface.co/datasets/albertge/MATH-500-gsm8k-format.Unroll-Qwen2.5-7B-Instruct_1754691065_eval_6a28_math500_top-5-voting_num_prune_attn_6
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754691065_eval_6a28_math500_top-5-voting_num_prune_attn_6
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 54.80%
Accuracy
Questions Solved
Total Questions
54.80%
274
500
OpenThoughts3-10k-dedup_cont3_without_math500math500-rubric-parser-math-verifypost-training-takehome-math500-bon16
MATH-500 Best-of-16 Post-Training Take-Home Results
A 50-problem study of test-time compute, based on the Hugging Face post-training take-home challenge. Nothing here trains or modifies a model: both the generator and the reward model stay frozen, and the only variable is how a final answer is chosen from 16 sampled candidates.
Construction
Filtered MATH-500 to levels 1-3, shuffled with seed 1, and selected 50 rows.
Generated one greedy solution per problem with… See the full description on the dataset page: https://huggingface.co/datasets/augustoFranke/post-training-takehome-math500-bon16.MATH500-ArCustom-OpenThinker-32B_1754029333_eval_466d_math500_prn_attn_7
chengfu0118/Custom-OpenThinker-32B_1754029333_eval_466d_math500_prn_attn_7
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 0.00%
Accuracy
Questions Solved
Total Questions
0.00%
0
500
Unroll-Qwen2.5-7B-Instruct_1754615103_eval_6a28_math500_skip_attn_idx_10
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754615103_eval_6a28_math500_skip_attn_idx_10
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 70.80%
Accuracy
Questions Solved
Total Questions
70.80%
354
500
Unroll-Qwen2.5-7B-Instruct_1754640953_eval_6a28_math500_skip_ffn_idx_1
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754640953_eval_6a28_math500_skip_ffn_idx_1
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 70.40%
Accuracy
Questions Solved
Total Questions
70.40%
352
500
Unroll-Qwen2.5-7B-Instruct_1754615061_eval_6a28_math500_skip_attn_idx_9
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754615061_eval_6a28_math500_skip_attn_idx_9
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 67.80%
Accuracy
Questions Solved
Total Questions
67.80%
339
500
Unroll-Qwen2.5-7B-Instruct_1754916227_eval_6a28_math500_geometric_num_prune_ffn_4_run-002
chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754916227_eval_6a28_math500_geometric_num_prune_ffn_4_run-002
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 38.20%
Accuracy
Questions Solved
Total Questions
38.20%
191
500
20250317-math500-sampling-solutions-32-tempmath500-rubric-math-verifyQwen2.5-7B-SFT-Math-Code-1M-1000-MATH500math500-output-audit
