datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool-reasoning-sft-RESEARCH-openresearcher-dataset-sft-deep-research-agent-data-cleaned
OpenResearcher Dataset - Cleaned & Restructured
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the OpenResearcher Dataset released by the TIGER-AI-Lab. The original dataset contains 96K+ long-horizon deep research trajectories generated by GPT-OSS-120B with native browser tools. This version converts the GPT-OSS channel-based message format into a standardized multi-turn tool-use conversation… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-openresearcher-dataset-sft-deep-research-agent-data-cleaned.verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified
Tool-Reasoning SFT — BrowseComp-Plus Runs (Cleaned & Rectified)
Multi-turn tool-use reasoning trajectories derived from grill-lab/browsecomp-plus-runs, converted to a structured SFT format following the interstellarninja/hermes_reasoning_tool_use convention.
Source
Based on the execution trajectories from "Revisiting Text Ranking in Deep Research" (arXiv:2602.21456):
Original data: grill-lab/browsecomp-plus-runs (MIT)
Format
Each row contains a messages… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified.tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified
Deep Research - Tulu SFT Data Cleaned Rectified
👥 Follow the Author
Supriti Vijay
Overview
This dataset is a cleaned and restructured version of the DR-TULU SFT dataset released by AllenAI's RL Research team. The original DR-TULU dataset represents significant work in creating high-quality training data for reasoning-enhanced language models with tool use capabilities. This version addresses structural issues in the original release while preserving… See the full description on the dataset page: https://huggingface.co/datasets/SupritiVijay/tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified.tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts
CodeScout Training Rollouts — Cleaned & Rectified
~40K multi-turn code localization agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Supports coupled (parallel) tool calls.
⚠️ Mid-training dataset. This dataset contains synthesized reasoning templates (not native chain-of-thought). It is suitable for mid-training to teach tool-use mechanics, FSM structure, and bash exploration patterns. It is not recommended as a final SFT… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts.tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source
Tool-Reasoning SFT — RLVR Retrieval Source Trajectories
156,381 multi-turn agentic retrieval trajectories across three document corpora, in a strict reasoning + tool-call format with validated FSM transitions. Each trajectory records a model searching a corpus, opening documents, and citing relevant passages to answer a question.
Author: Aman Priyanshu
Source Environments
Trajectories were collected against three RLVR retrieval environments from the FORMAT: Search -… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.tool-reasoning-sft-RESEARCH-explorations
Explorations Trajectories — Cleaned & Stripped
149,025 multi-turn code exploration agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions.
Origin
Derived from AmanPriyanshu/random-small-github-repositories and AmanPriyanshu/random-python-github-repositories.
Each trajectory is a search session where an agent navigates a GitHub repository using terminal commands to locate a target file. The agent reasons about project… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-explorations.kimi-k3-cyber-reasoning-distill
Kimi Cyber Reasoning
997 chain-of-thought records covering 13 cybersecurity disciplines and 4 systems engineering domains, distilled from the Kimi K3 reasoning model via API. Every record provides an explicit step-by-step <think> reasoning trace followed by a technical resolution, unified code diff fix, or structured tool invocation.
The dataset was curated as an anchor set for training, healing, and specializing compact reasoning models on systems security and tool calling… See the full description on the dataset page: https://huggingface.co/datasets/p-research/kimi-k3-cyber-reasoning-distill.Opus-4.6-Reasoning-2160x
Opus-4.6-Reasoning-2160x
2,160 high-quality reasoning traces generated by Claude Opus 4.6 via OpenRouter, covering mathematics, competitive programming, logic, science, and language tasks. Each example includes the full problem, an extended chain-of-thought, and a final solution — making the dataset suitable for supervised fine-tuning, chain-of-thought distillation, and reasoning-capability transfer to smaller models.
Originally generated as a batch of 3,305 examples; 1,145 were… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/Opus-4.6-Reasoning-2160x.MathChatSync-reasoning
MathChatSync-reasoning
The dataset was presented in the paper One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning.
The code for the paper can be found at: https://github.com/devrev/One-Pass-to-Reason
Dataset Description
MathChatSync-reasoning is a synthetic dialogue-based mathematics tutoring dataset designed to enable supervised training with explicit step-by-step reasoning. This dataset augments the… See the full description on the dataset page: https://huggingface.co/datasets/devrev-research/MathChatSync-reasoning.tool-reasoning-sft-RESEARCH-REDSearcher_SFT_10K
REDSearcher SFT 10K — Cleaned & Rectified
~8,850 multi-turn deep-search agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions.
Origin
Derived from Zchu/REDSearcher_SFT_10K.
REDSearcher is a deep search assistant dataset featuring rigorous, multi-step, multi-source investigations. Each trajectory contains a complex research question answered through iterative search → visit → reason → answer cycles, with extensive… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-REDSearcher_SFT_10K.tool-reasoning-sft-RESEARCH-OpenSeeker-v1-Data
OpenSeeker v1 — Cleaned & Rectified
7,189 multi-turn deep-search agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions.
Origin
Derived from OpenSeeker/OpenSeeker-v1-Data.
OpenSeeker is an open-source search agent system that democratizes access to frontier search capabilities by fully open-sourcing its training data. Fine-tuned on Qwen3-30B-A3B-Thinking-2507 with 11.7K training examples, it achieves state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-OpenSeeker-v1-Data.
