datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.glenans-isobars-archivearc-agi3-kimi-k2.7-ar25
ARC-AGI-3 ar25 — Agent Trajectories (kimi-k2.7)
Gameplay trajectories from the harness×model pair kimi-k2.7 playing the
ARC-AGI-3 game ar25, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ar25.arc_easyarc-agi3-agy-gemini3.1pro-tr87
ARC-AGI-3 tr87 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game tr87, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-tr87.arc_challengearc-agi3-agy-gemini3.1pro-g50t
ARC-AGI-3 g50t — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game g50t, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-g50t.arc-agi3-agy-gemini3.1pro-su15
ARC-AGI-3 su15 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game su15, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-su15.reward-projection-goal-generalisation-vlmTechnical-Architectures-Large
Technical Architectures Large (294k Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.ARC-Bench
ARC-Bench: An Open-Ended Autonomous-Research Benchmark Across Five Scientific Domains
The benchmark released with AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration.
ARC-Bench is a 55-topic, open-ended autonomous-research benchmark spanning
five scientific domains. Each topic is not a fixed-input/fixed-output task — it is a
research question plus a structured briefing. A research agent (or a human) must
take a topic from question →… See the full description on the dataset page: https://huggingface.co/datasets/AIMING-Lab-UNC/ARC-Bench.arc-steps
ARC Intermediate Solving Steps (arc-steps)
This dataset accompanies the paper TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning.
1,286,952 procedurally generated ARC-style records — 1,062,561 of them
with intermediate solving steps. Each record is an {input, steps, output} triple:
steps is a sequence of intermediate grids tracing a semantically meaningful solution
path from the input to the output, captured at human-annotated checkpoints of the
program that… See the full description on the dataset page: https://huggingface.co/datasets/lbn32/arc-steps.policystrategies-archive
🏛️ Open-Source Macro-Strategy, Financial History & Intelligence Archive
🌐 Overview & Institutional Mission
This public repository serves as the official open-source knowledge graph and metadata registry for r/policystrategies.
We aggregate, document, and cross-reference declassified historical intelligence dossiers, sovereign debt crises, systemic market manipulations, and geoeconomic conflicts using verified open-source intelligence (OSINT) and primary… See the full description on the dataset page: https://huggingface.co/datasets/stratigahq/policystrategies-archive.kepler-arc-agi-3-traces
Kepler 1.0 ARC-AGI-3 trace corpus
Run artifacts from Kepler 1.0, an open-source agent harness for the 25 public
ARC-AGI-3 games. A stock CLI coding agent
encodes its theory of each game as an executable world_model.py, certifies it
against the full recorded interaction history, plans inside the certified
model, and acts through a guarded channel that voids the plan on the first
misprediction.
Project page ·
Code ·
Paper ·
Integrity record
The canonical release contains two… See the full description on the dataset page: https://huggingface.co/datasets/cveinnt/kepler-arc-agi-3-traces.MAMIfreebsd-cvs-archive
📦 FreeBSD CVS Archive (C/C++)
Dataset Summary
FreeBSD CVS Archive (C/C++) is a large-scale dataset of source code extracted from the historical FreeBSD CVS repository. The dataset focuses on C and C++ source files, providing structured samples suitable for code modeling, analysis, and benchmarking.
Each sample includes:
the dataset source
commit year
extracted code content
token count (computed using GPT tokenizer)
This dataset is designed for:
code language modeling… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/freebsd-cvs-archive.archive-dolma3-pool-stratified
archive-dolma3-pool-stratified
ARCHIVE (pre-6T era): stratified sampling outputs from the 150B pool work.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_stratified
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-stratified.neurarch-architectures
Neurarch architecture corpus
36 neural network architectures kept as typed graphs rather than diagrams. Each entry carries every layer's type and parameters, the tensor shape propagated through it, an estimated parameter count, the verdict of 41 structural checks, and a link to the model.json an agent can fetch, edit and submit back for verification. Families span vision, language, recommendation, diffusion, biosignal and speech.
Size: 36 architectures
Licence: CC0 1.0… See the full description on the dataset page: https://huggingface.co/datasets/neurarch-ai/neurarch-architectures.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.udf-alpharl-out-archive
udf-alpharl-out-archive
Output archive for the paper "From Base Rollouts to RL Reasoning: A Budgeted Search Perspective" (https://arxiv.org/abs/2609.01274).
Code and configs: https://github.com/HALIS-sh/Searchlens_boptr
lm-eval-results-Kukedlc-Neural-4-ARC-7b-private
Dataset Card for Evaluation run of Kukedlc/Neural-4-ARC-7b
Dataset automatically created during the evaluation run of model Kukedlc/Neural-4-ARC-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-4-ARC-7b-private.lm-eval-results-arcee-ai-Clown-DPO-Extended-private
Dataset Card for Evaluation run of arcee-ai/Clown-DPO-Extended
Dataset automatically created during the evaluation run of model arcee-ai/Clown-DPO-Extended
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-arcee-ai-Clown-DPO-Extended-private.discord-archive
Discord Archive
This is an archive of messages from the Banodoco Discord community, where
technical and artistic practitioners have been discussing open source AI art for
the past three years.
The archive captures a long-running community record of people learning,
training, evaluating, and using open source AI art models in practice. It
contains discussion around model releases, workflows, tooling, troubleshooting,
creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.arch-selection-replay
Design selection replay: two trained campaigns and every selector tried
The records needed to score a design selector offline, in the shape of xRouteBench: every candidate design was trained once and the outcome kept, so a rule that picks which unexecuted design gets the GPU is evaluated by replay against the same rows as every previous rule, with no GPU and no model call. Two campaigns, 24 in-sample and 15 held-out designs, each row carrying what a selector may see (task… See the full description on the dataset page: https://huggingface.co/datasets/neurarch-ai/arch-selection-replay.arch-code-transfer-lpi-260903T0846-json-only-boundaryreview_arcade
ACL ARR Reviews - Review Arcade Project
This dataset contains paper reviews from the ACL ARR (Association for Computational Linguistics - Annual Review of Research) program. The reviews are organized into splits corresponding to papers that were accepted or rejected for publication, as well as specific subsets used for the research analysis.
This dataset was introduced in the paper Review Arcade: On the Human Alignment and Gameability of LLM Reviews.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/review_arcade.repro-conservation-laws-for-modern-neural-architectures-traces
Agent traces
Agent sessions published from a Trackio Logbook.
arc-synth-mcqa
ARC-Style Synthetic Science MCQs (teacher: Qwen3.8-27B)
Synthetic multiple-choice science questions generated for ARC-Challenge
fine-tuning, released for reproducibility of the companion model. Every file
that was used in training is included, along with the full audit trail.
Files
file
rows
what it is
clean_all.jsonl
6,861
v2 pool: generated, blind-label-verified, deduped, ARC-form-gated
clean_std.jsonl
4,630
non-negation subset of the above… See the full description on the dataset page: https://huggingface.co/datasets/minjujeon/arc-synth-mcqa.arcee-ai__Arcee-Nova-details
Dataset Card for Evaluation run of arcee-ai/Arcee-Nova
Dataset automatically created during the evaluation run of model arcee-ai/Arcee-Nova
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/arcee-ai__Arcee-Nova-details.
