datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.opensre
OpenSRE / OpenRCA dataset
Root cause analysis benchmark data: queries, incident records, telemetry (metrics, logs, traces), and query_alerts (per-row JSON derived from each query.csv).
Telemetry CSVs use different schemas by file type; load them by path (they are not merged into the Hub subset configs above).
Original archives are also described here: Google Drive.
Regenerating query_alerts
python3 scripts/query_csv_to_alert_json.py
trace-x-runtime-datanetopsbench-trace
NetOpsBench Agent Traces
This dataset contains sanitized NetOpsBench benchmark trace artifacts.
The legacy cross-model snapshot contains minimal-deepagent runs for
MiniMax M3, DeepSeek V4 Pro, Kimi K2.6, and OpenAI GPT-5.5 on the XS, Small,
Medium, and Large CLOS profiles.
The NetOpsBench v0.2 release adds a separately versioned
minimal-deepagent / deepseek-v4-pro snapshot across all seven built-in
profiles: XS, Small, Medium, Large, Xlarge, Fat-tree K=8, and Fat-tree K=12.
It… See the full description on the dataset page: https://huggingface.co/datasets/yyyyyt/netopsbench-trace.arc-agi-3-schema-traces-gpt56
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol
This release contains every gpt-5.6-sol gameplay trajectory produced on our
cluster with the world_model_v5 agent harness — 100 runs across the 25 public
ARC-AGI-3 games — plus a dependency-free scoring utility.
It is the GPT-5.6 Sol member of a family built by the same harness and the same
sanitizer, so trajectories can be compared game by game:
arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.ML-KEM-SideChannel-Traces
ML-KEM Side Channel Traces
Dataset Description
This dataset contains power traces captured from the Post-Quantum Cryptography (PQC) ML-KEM implementation of the PQM4[1] library (commit: a24bb4b), running on an STM32 Nucleo-L4R5ZI development board equipped with an ARM Cortex-M4 processor.
The traces were collected using a Rohde & Schwarz RTC1002 100 MHz digital oscilloscope. The purpose of this dataset is to evaluate the ML-KEM implementation for side-channel… See the full description on the dataset page: https://huggingface.co/datasets/ai-eldorado/ML-KEM-SideChannel-Traces.arc-agi-3-schema-traces-opus48
ARC-AGI-3 Schema Gameplay Trajectories — Claude Opus 4.8
This release contains the best claude-opus-4-8 / max trajectory for each of
the 25 public ARC-AGI-3 games, plus a dependency-free scoring utility. It is the
Opus 4.8 counterpart of
arc-agi-3-schema-traces-fable5,
produced by the same agent harness (world_model_v5) and the same sanitizer, so
the two collections can be compared game by game.
Each trajectory directory includes run.json, a streamed events.jsonl event
log… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-opus48.traceix-synthetic-classification-data
Traceix Synthetic Generated Data
Traceix Synthetic Generated Data is a synthetic tabular dataset for cybersecurity machine learning research, focused on static Windows PE-file metadata and binary classification workflows.
The dataset contains synthetically generated feature rows designed to resemble metadata patterns commonly extracted from Windows executable files during static analysis. It is intended for experimentation with malware/safe classification, anomaly detection, model… See the full description on the dataset page: https://huggingface.co/datasets/PerkinsFund/traceix-synthetic-classification-data.DDx-TRACE
DDx-TRACE
DDx-TRACE is a EuroRad-derived neuroradiology benchmark for evaluating multimodal differential-diagnosis reasoning and evidence tracing.
Files
data/eurorad_neuro_01_release.json: source nested manifest used by the benchmark code.
data/cases.csv: one row per case.
data/images.csv: one row per image/subfigure.
data/diagnostic_steps.csv: one row per diagnostic trajectory step.
data/imaging_examinations.csv: one row per imaging examination/evidence unit.… See the full description on the dataset page: https://huggingface.co/datasets/Anonym001/DDx-TRACE.Text2SQL_Workflow_Trace
Text2SQL Workflow Trace
Dataset Description
This dataset contains workflow traces for Text-to-SQL tasks, capturing the intermediate steps of translating natural language queries to executable SQL. It was used as input trace for the research presented in the paper:"HEXGEN-TEXT2SQL: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL Workflow" (arXiv:2505.05286).
The end-to-end Text-to-SQL queries collected in the dataset are from BIRD bench, and the trace… See the full description on the dataset page: https://huggingface.co/datasets/fredpeng/Text2SQL_Workflow_Trace.FM-Pochi-32B-Reasoning-TracesTRACE-it_CALAMITA
Dataset Card for TRACE-it Challenge @ CALAMITA 2024
TRACE-it (Testing Relative clAuses Comprehension through Entailment in ITalian) has been proposed as part of the CALAMITA Challenge, the special event dedicated to the evaluation of Large Language Models (LLMs) in Italian and co-located with the Tenth Italian Conference on Computational Linguistics (https://clic2024.ilc.cnr.it/calamita/).
The dataset focuses on evaluating LLM's understanding of a specific linguistic structure in… See the full description on the dataset page: https://huggingface.co/datasets/DominiqueBrunato/TRACE-it_CALAMITA.DDx-TRACE
DDx-TRACE
DDx-TRACE is a EuroRad-derived neuroradiology benchmark for evaluating multimodal differential-diagnosis reasoning and evidence tracing.
Files
data/eurorad_neuro_01_release.json: source nested manifest used by the benchmark code.
data/cases.csv: one row per case.
data/images.csv: one row per image/subfigure.
data/diagnostic_steps.csv: one row per diagnostic trajectory step.
data/imaging_examinations.csv: one row per imaging examination/evidence unit.… See the full description on the dataset page: https://huggingface.co/datasets/User3033/DDx-TRACE.npcsh-traces
enpisi-coder RL dataset
Judge-rated npcsh agent traces and derived RL training data for the
enpisi-coder model family.
Produced by scripts/rate_traces.py (LLM-as-judge) and
scripts/analyze_ratings.py; built into SFT/DPO/GRPO/PPO splits by
scripts/train_from_csv.py.
Splits
Split
Rows
Description
rated_traces
38388
Per-trace judge scores (correctness, tool_selection, efficiency, clarity, partial_credit, composite)
tasks
100
Benchmark task definitions… See the full description on the dataset page: https://huggingface.co/datasets/npc-worldwide/npcsh-traces.legal-causation-but-for-coherence-trace-v0.1Clarus Causation But-For Coherence Trace v0.1
This dataset tests whether a model can detect structural breakdown in legal causation.
It focuses on alignment between
defendant act
timeline
intervening events
harm outcome
counterfactual path
Causation is the backbone of liability.
When the chain breaks, the verdict eventually breaks.
This dataset measures that chain.
Core question
If the defendant act is removed, does the harm still occur.
If yes, the chain is incoherent.
If no, the chain holds.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/legal-causation-but-for-coherence-trace-v0.1.chatgpt_filtered_sft_traces_context_awaretrace-icml2026-repro-artifactscotempqa_for_sft_r1_tracessonnet_filtered_sft_traces_simplified_reasoninghighway_fast_v0_reasoning_traceschatgpt_filtered_sft_traces_simplified_reasoningsynthetic_cot_traces_cyphereval-agent-traceEval Agent Trace of a MLE Agent by Celestra. Sythetically Generated by gpt 5.2 thinking
gemini_filtered_sft_traces_simplified_reasoningsynthetic_cot_traces_clintoncotempqa_for_sft_4o_mini_summarized_r1_tracescotempqa_for_sft_r1_set_custom_incorrect_tracesMoE_expert_selection_traceCustom cl100k tokenized version of PubChem10M_SELFIES.
cotempqa_for_sft_r1_set_custom_traces
