kaetram
Datasets
All datasets matching “kaetram”kaetram-opd-2b
Kaetram OPD-2B — On-Policy Distillation Training Data
Training data for the Kaetram Qwen3.5-2B OPD models
(r1 ·
r2 ·
r3).
Each round is a distinct on-policy-distillation (OPD) dataset built from a 2B
agent's own gameplay rollouts, scored token-by-token against a stronger 4B teacher.
Round
Train records
Heldout
Init policy
round1
5,564
574
base Qwen3.5-2B
round2
7,024
825
merged r1
round3
8,856
1,040
merged r2
Configs
text (default, viewer) —… See the full description on the dataset page: https://huggingface.co/datasets/patnir41/kaetram-opd-2b.kaetram-qwen-rollouts
Kaetram Qwen Rollouts
Raw gameplay trajectories from Qwen3.5 agents (2B / 4B / 9B / 27B, plus the
2B OPD checkpoints) playing Kaetram via
a typed tool harness. These are the source logs behind the
Kaetram OPD-2B training
data and the
OPD model series.
Every config is Qwen self-play. Claude trajectories and any model fine-tuned on
them are deliberately excluded from this release.
Config
Model
Sessions
Turns
qwen9b-base
Qwen3.5-9B + scaffold
6,731
37,920
qwen27b-base… See the full description on the dataset page: https://huggingface.co/datasets/patnir41/kaetram-qwen-rollouts.
