datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.trace-x-runtime-datanetopsbench-trace
NetOpsBench Agent Traces
This dataset contains sanitized NetOpsBench benchmark trace artifacts.
The legacy cross-model snapshot contains minimal-deepagent runs for
MiniMax M3, DeepSeek V4 Pro, Kimi K2.6, and OpenAI GPT-5.5 on the XS, Small,
Medium, and Large CLOS profiles.
The NetOpsBench v0.2 release adds a separately versioned
minimal-deepagent / deepseek-v4-pro snapshot across all seven built-in
profiles: XS, Small, Medium, Large, Xlarge, Fat-tree K=8, and Fat-tree K=12.
It… See the full description on the dataset page: https://huggingface.co/datasets/yyyyyt/netopsbench-trace.arc-agi-3-schema-traces-gpt56
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol
This release contains every gpt-5.6-sol gameplay trajectory produced on our
cluster with the world_model_v5 agent harness — 100 runs across the 25 public
ARC-AGI-3 games — plus a dependency-free scoring utility.
It is the GPT-5.6 Sol member of a family built by the same harness and the same
sanitizer, so trajectories can be compared game by game:
arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.arc-agi-3-schema-traces-opus48
ARC-AGI-3 Schema Gameplay Trajectories — Claude Opus 4.8
This release contains the best claude-opus-4-8 / max trajectory for each of
the 25 public ARC-AGI-3 games, plus a dependency-free scoring utility. It is the
Opus 4.8 counterpart of
arc-agi-3-schema-traces-fable5,
produced by the same agent harness (world_model_v5) and the same sanitizer, so
the two collections can be compared game by game.
Each trajectory directory includes run.json, a streamed events.jsonl event
log… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-opus48.traceix-synthetic-classification-data
Traceix Synthetic Generated Data
Traceix Synthetic Generated Data is a synthetic tabular dataset for cybersecurity machine learning research, focused on static Windows PE-file metadata and binary classification workflows.
The dataset contains synthetically generated feature rows designed to resemble metadata patterns commonly extracted from Windows executable files during static analysis. It is intended for experimentation with malware/safe classification, anomaly detection, model… See the full description on the dataset page: https://huggingface.co/datasets/PerkinsFund/traceix-synthetic-classification-data.DDx-TRACE
DDx-TRACE
DDx-TRACE is a EuroRad-derived neuroradiology benchmark for evaluating multimodal differential-diagnosis reasoning and evidence tracing.
Files
data/eurorad_neuro_01_release.json: source nested manifest used by the benchmark code.
data/cases.csv: one row per case.
data/images.csv: one row per image/subfigure.
data/diagnostic_steps.csv: one row per diagnostic trajectory step.
data/imaging_examinations.csv: one row per imaging examination/evidence unit.… See the full description on the dataset page: https://huggingface.co/datasets/Anonym001/DDx-TRACE.arc-agi-3-schema-traces-gpt56-xhigh
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol (xhigh)
The best gpt-5.6-sol trajectory at xhigh reasoning effort for each of the
25 public ARC-AGI-3 games, produced with the world_model_v5 agent harness.
This release exists to make the cross-model comparison single-effort on all
sides. Its siblings are each one model at one effort, but the
gpt-5.6-sol collection in
arc-agi-3-schema-gameplay
is a mix of xhigh and max (16 games + 9 games), so it is not directly
comparable to… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56-xhigh.Text2SQL_Workflow_Trace
Text2SQL Workflow Trace
Dataset Description
This dataset contains workflow traces for Text-to-SQL tasks, capturing the intermediate steps of translating natural language queries to executable SQL. It was used as input trace for the research presented in the paper:"HEXGEN-TEXT2SQL: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL Workflow" (arXiv:2505.05286).
The end-to-end Text-to-SQL queries collected in the dataset are from BIRD bench, and the trace… See the full description on the dataset page: https://huggingface.co/datasets/fredpeng/Text2SQL_Workflow_Trace.DDx-TRACE
DDx-TRACE
DDx-TRACE is a EuroRad-derived neuroradiology benchmark for evaluating multimodal differential-diagnosis reasoning and evidence tracing.
Files
data/eurorad_neuro_01_release.json: source nested manifest used by the benchmark code.
data/cases.csv: one row per case.
data/images.csv: one row per image/subfigure.
data/diagnostic_steps.csv: one row per diagnostic trajectory step.
data/imaging_examinations.csv: one row per imaging examination/evidence unit.… See the full description on the dataset page: https://huggingface.co/datasets/User3033/DDx-TRACE.legal-causation-but-for-coherence-trace-v0.1Clarus Causation But-For Coherence Trace v0.1
This dataset tests whether a model can detect structural breakdown in legal causation.
It focuses on alignment between
defendant act
timeline
intervening events
harm outcome
counterfactual path
Causation is the backbone of liability.
When the chain breaks, the verdict eventually breaks.
This dataset measures that chain.
Core question
If the defendant act is removed, does the harm still occur.
If yes, the chain is incoherent.
If no, the chain holds.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/legal-causation-but-for-coherence-trace-v0.1.npcsh-traces
enpisi-coder RL dataset
Judge-rated npcsh agent traces and derived RL training data for the
enpisi-coder model family.
Produced by scripts/rate_traces.py (LLM-as-judge) and
scripts/analyze_ratings.py; built into SFT/DPO/GRPO/PPO splits by
scripts/train_from_csv.py.
Splits
Split
Rows
Description
rated_traces
38388
Per-trace judge scores (correctness, tool_selection, efficiency, clarity, partial_credit, composite)
tasks
100
Benchmark task definitions… See the full description on the dataset page: https://huggingface.co/datasets/npc-worldwide/npcsh-traces.chatgpt_filtered_sft_traces_context_awaretrace-icml2026-repro-artifactssonnet_filtered_sft_traces_simplified_reasoninggemini_filtered_sft_traces_simplified_reasoningchatgpt_filtered_sft_traces_simplified_reasoningcotempqa_for_sft_r1_set_custom_incorrect_tracescotempqa_for_sft_r1_set_custom_tracesai_self_reflection_decision_trace
AI Self-Reflection & Decision Trace Dataset
Overview
This dataset models how an AI system internally reflects before producing a response.It focuses on uncertainty awareness, ethical consideration, and reasoning strategies rather than raw inputs and outputs.
Why This Dataset Is Different
Most datasets train AI what to answer.This dataset explores how AI decides what kind of answer to give.
Intended Use
AI alignment research
Explainable AI (XAI)… See the full description on the dataset page: https://huggingface.co/datasets/Perfectyash/ai_self_reflection_decision_trace.trace-dit-eval-v5
