datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.arm_O3
REBench — arm / O3
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O3.gr1_arms_waist-CuttingboardToPanarm_O0
REBench — arm / O0
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O0.gr1_arms_waist-CuttingboardToCardboardBoxgr1_arms_waist-PlaceMilkToMicrowavegr1_arms_waist-WineToCabinetgr1_arms_waist-PlacematToBasketgpt-5.5-agentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
gpt 5.5 Agent Traces
This directory contains raw agent trace files generated by teich. (I also dropped in some of my own personal traces)
All assistant responses were generated by openai/gpt-5.5.
JSONL files: 88
Training-ready tools
A complete configured tools schema snapshot is embedded in the… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/gpt-5.5-agent.gr1_arms_waist-TrayToTieredShelfgr1_arms_waist-TrayToPlategpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Kimi K2.6 Claude Code Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by moonshotai/kimi-k2.6.
JSONL files: 36
Format
Each file is newline-delimited JSON representing a single captured agent session.
The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.qwen3.7-max-pi-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Qwen3.7 Max Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by qwen/qwen3.7-max.
JSONL files: 47
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.
Use it… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-pi-traces.armnet-demo-leaderboardarm_O1
REBench — arm / O1
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O1.gr1_arms_waist-TrayToPotarm_O2
REBench — arm / O2
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O2.gr1_arms_waist-PlateToCardboardBoxAgiBot-g1_robotic_arm_picks_up_battery
AgiBot-g1_robotic_arm_picks_up_battery
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
place
pick
grasp
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_robotic_arm_picks_up_battery.AgiBot-g1_robotic_arm_picks_up_parts
AgiBot-g1_robotic_arm_picks_up_parts
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
place
pick
grasp
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_robotic_arm_picks_up_parts.minimax-m3-claude-code-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Minimax M3 Claude Code Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by minimax/minimax-m3.
JSONL files: 31
Format
Each file is newline-delimited JSON representing a single captured agent session.
The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m3-claude-code-traces.phylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.dexmg-two-arm-drawer-cleanupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "robomimic",
"total_episodes": 1026,
"total_frames": 298235,
"total_tasks": 1,
"total_videos": 3078,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1026"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/dexmg-two-arm-drawer-cleanup.gr1_arms_waist-PlateToPanqrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
qrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
icon_arm-rolloutsthird_arm_01This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 10457,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/acrampette/third_arm_01.gr1_arms_waist-PlaceBottleToCabinetqrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offline-armorm
qrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offline-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
