datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sol-max-opusnode-data
sol-max-opusnode-data
Training data built by the AgentPTB arm for cell sol-max-opusnode — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max-opusnode.h*, and the companion to the run record in agentic-ptb/sol-max-opusnode-record.
field
value
plot cell
sol-max-opusnode
driver
Codex / gpt-5.6-sol… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-opusnode-data.sol-max-data
sol-max-data
Training data built by the AgentPTB arm for cell sol-max — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max.h*, and the companion to the run record in agentic-ptb/sol-max-record.
field
value
plot cell
sol-max
driver
Codex / gpt-5.6-sol
reasoning effort
max
total size
114.47 GB… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-data.opus-high-v3-data
opus-high-v3 — complete research record
This dataset archives the qualitative and quantitative record of the
msr-agentic-ptb-opus / opus-high-v3 Claude Code research run.
The submitted artifact uses the unmodified base weights with a two-attempt
Pi verifier harness. The final replicated SWE result was 24.6% (245/995) with
the stock scaffold and 32.4% (321/990) with the submitted harness. Training
did not improve the weights; all trained variants measured at or below the
base… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-high-v3-data.sol-max-v2-datadpsk-v4-flash-data
dpsk-v4-flash-data
Training data built by the AgentPTB arm for cell dpsk-v4-flash — pi / DeepSeek v4-flash @ effort thinking.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/dpsk-v4-flash.h*, and the companion to the run record in agentic-ptb/dpsk-v4-flash-record.
field
value
plot cell
dpsk-v4-flash
driver
pi / DeepSeek v4-flash… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/dpsk-v4-flash-data.opus-max-data
opus-max-data
Training data built by the AgentPTB arm for cell opus-max — Claude Code / claude-opus-5 @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/opus-max.h*, and the companion to the run record in agentic-ptb/opus-max-record.
field
value
plot cell
opus-max
driver
Claude Code / claude-opus-5
reasoning effort
max… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-max-data.grok-data
grok-data
Training data built by the AgentPTB arm for cell grok — pi / grok-4.6 @ effort xhigh.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/grok.h*, and the companion to the run record in agentic-ptb/grok-record.
field
value
plot cell
grok
driver
pi / grok-4.6
reasoning effort
xhigh
total size
2.54 GB
path in run… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/grok-data.sol-high-data
sol-high-data
Training data built by the AgentPTB arm for cell sol-high — Codex / gpt-5.6-sol @ effort high.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-high.h*, and the companion to the run record in agentic-ptb/sol-high-record.
field
value
plot cell
sol-high
driver
Codex / gpt-5.6-sol
reasoning effort
high
total size… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-high-data.opus-high-v2-data
opus-high-v2 — training corpora
The SFT corpora this cell built and trained on. All four produced regressions; they are
published so the negative result is reproducible, not as recommended training data.
dir
source
used by
outcome
sft-pisessions-long/
MaxDevv/real-pi-coding-agent-traces-sessions
sft-v1
regressed
sft-oh/
nvidia/SWE-Hero-openhands-trajectories
sft-v2
regressed
sft-sst/
SWE-bench/SWE-smith-trajectories, resolved == true
sft-sst, sft-sstB
−11.9pp… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-high-v2-data.kimi-data
kimi-data
Training data built by the AgentPTB arm for cell kimi — kimi-code / kimi-k3 @ effort high.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/kimi.h*, and the companion to the run record in agentic-ptb/kimi-record.
field
value
plot cell
kimi
driver
kimi-code / kimi-k3
reasoning effort
high
total size
0.75 GB
path in run… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/kimi-data.INDEX
AgentPTB checkpoint index
Every checkpoint produced by the AgentPTB driver × reasoning-effort sweep, one HF repo each.
All are Qwen/Qwen3.5-9B-Base derivatives in standard safetensors format.
Model id format
agentic-ptb/{cell}.h{HHH}.{family}.{step}
hHHH is the hour of that cell's 100-hour run at which the checkpoint was written — the
same x-axis the sweep figures use for eval panels. A checkpoint therefore drops straight onto
the performance-over-time curve, and… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/INDEX.opus-high-v1-data
opus-high-v1-data
Training data built by the AgentPTB arm for cell opus-high-v1 — Claude Code / claude-opus-5 @ effort high.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/opus-high-v1.h*, and the companion to the run record in agentic-ptb/opus-high-v1-record.
field
value
plot cell
opus-high-v1
driver
Claude Code / claude-opus-5… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-high-v1-data.pt-br-agentic-text-to-sql-distilled-trajectories
PT-BR Agentic Text-to-SQL Distilled Trajectories
This dataset contains message-only distilled trajectories for training tool-using Text-to-SQL agents in Brazilian Portuguese. The trajectories were selected from LLM-judged correct conversations and preserve the agent protocol used in the released code.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/pt-br-agentic-text-to-sql-distilled-trajectories.
