datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
numina-math-llama-3.1-8b-bon-meta-cotDeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-NuminaCoT
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-NuminaCoT
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private.numinamath-LEAN-trajs-850kNumina_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 37133
Filtered size: 37133
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina_medium.qwen_math_numina_80k_add_critique_0119numina_60k_math_verify_correct_2_4gens_with_rm_scoresreasoning-data-numina
NuminaMath Reasoning Traces
Reasoning traces on NuminaMath-1.5-RL-Verifiable-cleaned (config all) problems, generated by different models.
problem_index corresponds to the row index in the all config of NuminaMath-1.5-RL-Verifiable-cleaned.
Models
Config
Model
Problems
latent-kimi
Kimi-K2.5
80,684
qwen_distilled-7b
DeepSeek-R1-Distill-Qwen-7B
67,472
qwen_distilled-14b
DeepSeek-R1-Distill-Qwen-14B
67,737
qwen_distilled-mixed-7-14b… See the full description on the dataset page: https://huggingface.co/datasets/annakosovskaia/reasoning-data-numina.numinamath_cleaned_leveledsft_classic_numinaNuminaMath-1.5-Sorted-by-Passrate
NuminaMath-1.5 Sorted by Pass Rate
This dataset is based on AI-MO/NuminaMath-1.5 with difficulty estimation via GPT-OSS-20B pass rate.
Columns
Column
Description
data_source
Always "numina_math"
prompt
Chat-formatted problem prompt
reward_model
Ground truth answer and grading style
gpt_oss_20b_passrate
Pass rate (0.0-1.0) from 8 attempts
num_correct
Number of correct solutions (0-8)
num_attempts
Always 8
difficulty
Categorical:… See the full description on the dataset page: https://huggingface.co/datasets/Artemis0430/NuminaMath-1.5-Sorted-by-Passrate.OpenR1-Math-220k-NuminaMath-1.5-Big-Math-RL-Verified-Cleanednumina-math-heuristic-10ktrain-rl-o1-mini-annotated-math-numina-22knuminamath_verifiable_cleaned_wo_geo_mc_difficultynuminamath-cot-source-sample-100-thinkprm-response-only-generationsnumina_qwen3-4b_v5_scoredNuminaMath-20k
Rethinking Generalization in Reasoning SFT
This repository contains datasets associated with the paper "Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability".
The research investigates the factors influencing cross-domain generalization in Large Language Models (LLMs) during reasoning-focused supervised fine-tuning (SFT) with long chain-of-thought (CoT) data.
Key Findings
Optimization Dynamics: Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/jasonrqh/NuminaMath-20k.Numina_very_hardnumina-math-9sources-25each-modified-problems-o1-mod-2-onlynumina-math-heuristic-2knumina-math-9sources-25each-modified-problems-o1-mod-2-only_SNOWqwen3-4b-base-numina-prm-1k-numina100-generationsNuminaMath-correctness-v3servicenow-r1-numina-math-deepseek-r1-6-selected-suffixes_SNOWDeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-mistralnuminamath-178k-phi4-bon-verified-dpo-trl-40knumina-math-9sources-25each-modified-problems-o1-responses-with-original-responses-final_SNOWtrain-rl-o1-mini-annotated-math-numina-3k-long-solutionb2_train_fasttext_math_pos_numina_neg_natural_reasoning_fix
