verified-reasoning
qwen-7b-fsdp-magpie-reasoning-v1-10k-verified-onlyqwen-3b-fsdp-magpie-reasoning-v1-10k-verified-only-checkpoint672qwen-1_5b-fsdp-magpie-reasoning-v1-10k-verified-onlyqwen-3b-fsdp-magpie-reasoning-v1-10k-verified-onlyqwen-1_5b-fsdp-magpie-reasoning-v1-10k-verified-only-checkpoint1120qwen-3b-fsdp-magpie-reasoning-v1-10k-verified-only-checkpoint1344qwen-3b-fsdp-magpie-reasoning-v1-10k-verified-only-checkpoint896qwen-05b-fsdp-magpie-reasoning-v1-10k-verified-only
SWE-Bench-Verified-O1-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.verified-reasoning-o1-gpqa-mmlu-pro
Reasoning PRM Preference Dataset
This dataset contains reasoning traces from multiple sources (GPQA Diamond and MMLU Pro), labeled with preference information based on correctness verification.
Dataset Description
Overview
The dataset consists of reasoning problems and their solutions, where each example has been verified for correctness and labeled with a preference score. It combines data from two main sources:
GPQA Diamond
MMLU Pro
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ariaattarml/verified-reasoning-o1-gpqa-mmlu-pro.chart-reasoning-verified
chart-reasoning-verified
Chart reasoning examples generated from an explicit latent representation.
The data, the question and the answer are computed before the chart is
drawn, so the image is a rendering of known ground truth rather than the
source of it. No model was asked to label anything.
Each row carries both a rendered chart and a text serialisation of the same
chart, so the set is usable for vision-language training and for text-only
language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.verified-math-reasoning
verified-math-reasoning (CargoDash flagship recipe)
A CargoDash framework demonstration. 999-row showcase of three-layer,
program-verified, vote-stratified math reasoning traces — the dataset is
small on purpose (its job is to prove the framework works on real
production LLM endpoints, not to be a serious math benchmark). Each row
carries three independent chain-of-thought solutions to the same problem
(from DeepSeek, Doubao, and Qwen3.5) plus a programmatically extracted
\boxed{}… See the full description on the dataset page: https://huggingface.co/datasets/Bdyskov/verified-math-reasoning.
