datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
hermes-flight-recorder-self-improving-agent-trajectories
Hermes Flight Recorder Self-Improving Agent Trajectories
This public-safe synthetic dataset contains 800 governed agent trajectories
for supervised tool-use training, 120 development tasks, and a separately
frozen set of 150 final evaluation tasks. It demonstrates how recorded
successful executions and reviewed safety
refusals can become training data without publishing user traces.
Files
train_trajectories.jsonl: 800 training-only conversational tool-use rows… See the full description on the dataset page: https://huggingface.co/datasets/zwright/hermes-flight-recorder-self-improving-agent-trajectories.self-improving-preferences-sftsdtself_improving_old2self-improving-preferences-sftsd1self-improving-preferences-sftsd0self-improving-preferences-sftsd2self_improving_oldself_improvingSI2CA-Training-TrajectoriesDataset Card for SI2CA-Training-Trajectories
[🌐 Website] •
[🤗 Dataset] •
[📜 Paper] •
[🐱 GitHub]
💡 Introduction
This dataset consists of 32,340 coding-agent trajectories generated by Qwen3.5-122B-A10B on the same 10,780 executable Python SWE tasks under the three trajectory-curation settings of Section 4.4 of the paper: standard sampling, full self-judgement, and an efficient discovered strategy found by the recursive self-improvement framework. Each task is… See the full description on the dataset page: https://huggingface.co/datasets/Self-Improving-Coding-Agents/SI2CA-Training-Trajectories.
