CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LeoZotos /math500tabularn<1K0 likes715 downloads1mo agoHugging Face02alessiodevoto /math500tabularn<1K0 likes657 downloads2mo agoHugging Face03ChrisMcCormick /math500-cot-deepseek-r1-1.5b MATH-500 CoT completions (DeepSeek-R1-Distill-Qwen-1.5B) Successful chain-of-thought completions for HuggingFaceH4/MATH-500 test problems, generated with deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B via vLLM. Files File Description records.parquet Main dataset: correct completions as token IDs manifest.json Schema, tokenizer, run ids, decoding config problem_index.json unique_id → problem_idx in MATH-500 test subject_max_tokens.json Per-subject completion… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/math500-cot-deepseek-r1-1.5b.tabulartext-generation1K<n<10K0 likes72 downloads4mo agoHugging Face04lulululuyi /R-HORIZON-Math500tabular1K<n<10K0 likes69 downloads1y agoHugging Face05meituan-longcat /R-HORIZON-Math500 R-HORIZON How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 📃 Paper • 🌐 Project Page • 🤗 Dataset R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Math500.tabular1K<n<10K1 likes67 downloads11mo agoHugging Face06sibasmarakp /Qwen2.5-1.5B-Instruct-uPRM-70B-T80-math500-best_of_n-completionstabular1K<n<10K0 likes61 downloads24d agoHugging Face07sibasmarakp /Llama-3.2-1B-Instruct-uPRM-32B-T80-math500-best_of_n-completionstabular1K<n<10K0 likes54 downloads24d agoHugging Face08sibasmarakp /Llama-3.2-1B-Instruct-uPRM-70B-T80-math500-best_of_n-completionstabular1K<n<10K0 likes53 downloads24d agoHugging Face09sibasmarakp /Qwen2.5-1.5B-Instruct-uPRM-32B-T80-math500-best_of_n-completionstabular1K<n<10K0 likes47 downloads24d agoHugging Face10amphora /m-math500tabularn<1K0 likes40 downloads2y agoHugging Face11osieosie /mixed_sft_math500_128_s1_tulu2_sft_s1_1.0pcttabular10K<n<100K0 likes37 downloads1y agoHugging Face12chengfu0118 /Custom-Bespoke-Stratos-7B_1753961306_eval_466d_math500_skip_attn_1 chengfu0118/Custom-Bespoke-Stratos-7B_1753961306_eval_466d_math500_skip_attn_1 Precomputed model outputs for evaluation. Evaluation Results MATH500 Accuracy: 1.80% Accuracy Questions Solved Total Questions 1.80% 9 500 tabularn<1K0 likes31 downloads1y agoHugging Face13cmpatino /math500-bon-weighted-results MATH-500 Best-of-N Weighted Selection Results Dataset Description This dataset contains the results of evaluating Best-of-N weighted selection on a subset of the MATH-500 benchmark. It was created as part of a HuggingFace internship exercise exploring how test-time compute scaling with reward models can improve LLM performance on math problems. How It Was Constructed 1. Problem Selection Started from the HuggingFaceH4/MATH-500 dataset (500 problems)… See the full description on the dataset page: https://huggingface.co/datasets/cmpatino/math500-bon-weighted-results.tabulartext-generationn<1K1 likes31 downloads5mo agoHugging Face14idoazou /math500-typostabular10K<n<100K0 likes31 downloads2mo agoHugging Face15mikasenghaas /math500-checktabularn<1K0 likes28 downloads1y agoHugging Face16albertge /MATH-500-gsm8k-format MATH-500-gsm8k-format Dataset Description This dataset contains 500 mathematical problems from the MATH-500 benchmark, converted to GSM8K format for step-by-step reasoning. Dataset Summary Source: HuggingFaceH4/MATH-500 Format: GSM8K-style step-by-step solutions with inline computation annotations Size: 500 problems Split: Test (original MATH-500 test split) Conversion Process The original MATH-500 solutions (which use LaTeX notation and… See the full description on the dataset page: https://huggingface.co/datasets/albertge/MATH-500-gsm8k-format.tabulartext-generationn<1K0 likes28 downloads11mo agoHugging Face17chengfu0118 /Unroll-Qwen2.5-7B-Instruct_1754691065_eval_6a28_math500_top-5-voting_num_prune_attn_6 chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754691065_eval_6a28_math500_top-5-voting_num_prune_attn_6 Precomputed model outputs for evaluation. Evaluation Results MATH500 Accuracy: 54.80% Accuracy Questions Solved Total Questions 54.80% 274 500 tabularn<1K0 likes26 downloads1y agoHugging Face18reasoningMIA /OpenThoughts3-10k-dedup_cont3_without_math500tabular10K<n<100K0 likes24 downloads1y agoHugging Face19mikasenghaas /math500-rubric-parser-math-verifytabularn<1K0 likes24 downloads1y agoHugging Face20augustoFranke /post-training-takehome-math500-bon16 MATH-500 Best-of-16 Post-Training Take-Home Results A 50-problem study of test-time compute, based on the Hugging Face post-training take-home challenge. Nothing here trains or modifies a model: both the generator and the reward model stay frozen, and the only variable is how a final answer is chosen from 16 sampled candidates. Construction Filtered MATH-500 to levels 1-3, shuffled with seed 1, and selected 50 rows. Generated one greedy solution per problem with… See the full description on the dataset page: https://huggingface.co/datasets/augustoFranke/post-training-takehome-math500-bon16.tabulartext-generationn<1K0 likes24 downloads2mo agoHugging Face21riturajj /MATH500-Artabularn<1K0 likes23 downloads1y agoHugging Face22chengfu0118 /Custom-OpenThinker-32B_1754029333_eval_466d_math500_prn_attn_7 chengfu0118/Custom-OpenThinker-32B_1754029333_eval_466d_math500_prn_attn_7 Precomputed model outputs for evaluation. Evaluation Results MATH500 Accuracy: 0.00% Accuracy Questions Solved Total Questions 0.00% 0 500 tabularn<1K0 likes22 downloads1y agoHugging Face23chengfu0118 /Unroll-Qwen2.5-7B-Instruct_1754615103_eval_6a28_math500_skip_attn_idx_10 chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754615103_eval_6a28_math500_skip_attn_idx_10 Precomputed model outputs for evaluation. Evaluation Results MATH500 Accuracy: 70.80% Accuracy Questions Solved Total Questions 70.80% 354 500 tabularn<1K0 likes22 downloads1y agoHugging Face24chengfu0118 /Unroll-Qwen2.5-7B-Instruct_1754640953_eval_6a28_math500_skip_ffn_idx_1 chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754640953_eval_6a28_math500_skip_ffn_idx_1 Precomputed model outputs for evaluation. Evaluation Results MATH500 Accuracy: 70.40% Accuracy Questions Solved Total Questions 70.40% 352 500 tabularn<1K0 likes21 downloads1y agoHugging Face25chengfu0118 /Unroll-Qwen2.5-7B-Instruct_1754615061_eval_6a28_math500_skip_attn_idx_9 chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754615061_eval_6a28_math500_skip_attn_idx_9 Precomputed model outputs for evaluation. Evaluation Results MATH500 Accuracy: 67.80% Accuracy Questions Solved Total Questions 67.80% 339 500 tabularn<1K0 likes19 downloads1y agoHugging Face26chengfu0118 /Unroll-Qwen2.5-7B-Instruct_1754916227_eval_6a28_math500_geometric_num_prune_ffn_4_run-002 chengfu0118/Unroll-Qwen2.5-7B-Instruct_1754916227_eval_6a28_math500_geometric_num_prune_ffn_4_run-002 Precomputed model outputs for evaluation. Evaluation Results MATH500 Accuracy: 38.20% Accuracy Questions Solved Total Questions 38.20% 191 500 tabularn<1K0 likes19 downloads1y agoHugging Face27hzy /20250317-math500-sampling-solutions-32-temptabular1K<n<10K0 likes18 downloads2y agoHugging Face28mikasenghaas /math500-rubric-math-verifytabularn<1K0 likes18 downloads1y agoHugging Face29mikasenghaas /Qwen2.5-7B-SFT-Math-Code-1M-1000-MATH500tabularn<1K0 likes18 downloads1y agoHugging Face30vibhuiitj /math500-output-audittabularn<1K0 likes18 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.