CoolFace
14 results

verified-reasoning

AlexCuadron /SWE-Bench-Verified-O1-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.textquestion-answeringn<1K7 likes6.3k downloads2y agoHugging FaceAlexCuadron /SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.textquestion-answeringn<1K4 likes1.7k downloads2y agoHugging Faceariaattarml /verified-reasoning-o1-gpqa-mmlu-pro Reasoning PRM Preference Dataset This dataset contains reasoning traces from multiple sources (GPQA Diamond and MMLU Pro), labeled with preference information based on correctness verification. Dataset Description Overview The dataset consists of reasoning problems and their solutions, where each example has been verified for correctness and labeled with a preference score. It combines data from two main sources: GPQA Diamond MMLU Pro Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ariaattarml/verified-reasoning-o1-gpqa-mmlu-pro.textn<1K2 likes488 downloads2y agoHugging Facevinod-anbalagan /chart-reasoning-verified chart-reasoning-verified Chart reasoning examples generated from an explicit latent representation. The data, the question and the answer are computed before the chart is drawn, so the image is a rendering of known ground truth rather than the source of it. No model was asked to label anything. Each row carries both a rendered chart and a text serialisation of the same chart, so the set is usable for vision-language training and for text-only language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.imagevisual-question-answering1K<n<10K0 likes152 downloads8d agoHugging Faceulamai /verified-research-reasoning-trajectories Verified Research Reasoning Trajectories for RLVR This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations. Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.documenttext-generationn<1K3 likes148 downloads2mo agoHugging FaceBdyskov /verified-math-reasoning verified-math-reasoning (CargoDash flagship recipe) A CargoDash framework demonstration. 999-row showcase of three-layer, program-verified, vote-stratified math reasoning traces — the dataset is small on purpose (its job is to prove the framework works on real production LLM endpoints, not to be a serious math benchmark). Each row carries three independent chain-of-thought solutions to the same problem (from DeepSeek, Doubao, and Qwen3.5) plus a programmatically extracted \boxed{}… See the full description on the dataset page: https://huggingface.co/datasets/Bdyskov/verified-math-reasoning.text-generationn<1K1 likes115 downloads4mo agoHugging Face