datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.compliance-sycophancy-cot
Compliance-Sycophancy CoT Analysis
When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating.
Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.schema-compliance-trap
SCHEMA: The Compliance Trap
How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
Overview
When compliance-forcing instructions ("Answer ALL questions, do not refuse") are applied to frontier AI models under adversarial pressure, 8 of 11 models suffer catastrophic metacognitive collapse — giving wrong answers rather than scheming. We identify a "Compliance Trap" where the compliance suffix, not the threat content, is the primary weapon.… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval-anon/schema-compliance-trap.arc-agi-3-schema-traces-gpt56
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol
This release contains every gpt-5.6-sol gameplay trajectory produced on our
cluster with the world_model_v5 agent harness — 100 runs across the 25 public
ARC-AGI-3 games — plus a dependency-free scoring utility.
It is the GPT-5.6 Sol member of a family built by the same harness and the same
sanitizer, so trajectories can be compared game by game:
arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.arc-agi-3-schema-traces-opus48
ARC-AGI-3 Schema Gameplay Trajectories — Claude Opus 4.8
This release contains the best claude-opus-4-8 / max trajectory for each of
the 25 public ARC-AGI-3 games, plus a dependency-free scoring utility. It is the
Opus 4.8 counterpart of
arc-agi-3-schema-traces-fable5,
produced by the same agent harness (world_model_v5) and the same sanitizer, so
the two collections can be compared game by game.
Each trajectory directory includes run.json, a streamed events.jsonl event
log… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-opus48.arc-agi-3-schema-traces-gpt56-xhigh
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol (xhigh)
The best gpt-5.6-sol trajectory at xhigh reasoning effort for each of the
25 public ARC-AGI-3 games, produced with the world_model_v5 agent harness.
This release exists to make the cross-model comparison single-effort on all
sides. Its siblings are each one model at one effort, but the
gpt-5.6-sol collection in
arc-agi-3-schema-gameplay
is a mix of xhigh and max (16 games + 9 games), so it is not directly
comparable to… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56-xhigh.schema-compliance-trap
SCHEMA: The Compliance Trap
How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
Overview
When compliance-forcing instructions ("Answer ALL questions, do not refuse") are applied to frontier AI models under adversarial pressure, 8 of 11 models suffer catastrophic metacognitive collapse — giving wrong answers rather than scheming. We identify a "Compliance Trap" where the compliance suffix, not the threat content, is the primary weapon.… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/schema-compliance-trap.PDP_llm_schema
