datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-agi3-agy-gemini3.1pro-tr87
ARC-AGI-3 tr87 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game tr87, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-tr87.arc-agi3-agy-gemini3.1pro-ar25
ARC-AGI-3 ar25 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game ar25, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-ar25.arc-agi3-agy-gemini3.1pro-g50t
ARC-AGI-3 g50t — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game g50t, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-g50t.discoverphysics-gemini3.1-pro-ara
DiscoverPhysics × Gemini 3.1 Pro High (Antigravity CLI) — 11-world ARA knowledge artifacts
Agent-Native Research Artifacts (ARA) produced by a Gemini 3.1 Pro High (Antigravity CLI) coding-agent session solving all
11 worlds of the DiscoverPhysics scientific-discovery
benchmark (seed 0, noise_frac 0.075, ≤16 experiment rounds), driven through the same
harness-agnostic bridge and ARA scaffold as the sibling fable run. Official verdicts:
1/11 PASS — criteria and per-world numbers… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/discoverphysics-gemini3.1-pro-ara.arc-agi3-agy-gemini3.1pro-s5i5
ARC-AGI-3 s5i5 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game s5i5, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-s5i5.arc-agi3-agy-gemini3.1pro-ls20
ARC-AGI-3 ls20 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game ls20, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-ls20.arc-agi3-agy-gemini3.1pro-su15
ARC-AGI-3 su15 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game su15, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-su15.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.arc-agi3-agy-gemini3.1pro-r11l
ARC-AGI-3 r11l — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game r11l, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-r11l.gemini-3.1-pro-2048-reasoning-1100xlhtb-gemini3.1-pro-high-trajectories
LHTB trajectories — gemini3.1-pro-high
Detailed agent trajectories on Long-Horizon Terminal-Bench
(46 long-horizon terminal tasks, hidden-verifier dense rewards).
Harness: LHTB's continue-until-timeout patched Harbor,
90-min budget per task (override_timeout_sec: 5400), n_attempts: 1.
Agent: installed CLI agent (agy), NOT the paper's Terminus-2 —
same tasks/budget/harness, different agent scaffold.
Per trial: agent/trajectory.json (ATIF v1.5: steps, tool calls, reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/lhtb-gemini3.1-pro-high-trajectories.Gemini-3.1-Pro-GLM5-CharactersPrompts generated by Gemini 3.1 Pro.
Responses generated by Gemini 3.1 Pro.
Reasoning traces:
Step 1: Generated by GLM 5 which was provided the original system prompt / knowledge
Step 2: Edited by Gemini to fix any contradictions with the existing response
Step 3: Edited again to remove / reduce drafting and remove reasoning related to safety / refusals
System prompt was generated based on the original and whatever constraints / rules etc. had been mentioned in the reasoning trace.
A small… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Gemini-3.1-Pro-GLM5-Characters.arc-agi3-agy-gemini3.1pro-ft09
ARC-AGI-3 ft09 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game ft09, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-ft09.Gemini-3.1-Pro-SmallWikiWas messing around with Gemini 3.1 Pro and quite liked its knowledge.
Was using it to generate some anime related stuff as usual, but decided to dump it some wiki articles after I was done with that and also generate this more generalized dataset too.
Each wiki page has 5 samples generated.
Prompts were generated by Gemini 3.1 Pro, responses were generated by Gemini 3.1 Pro, which was provided the wikipage. Random instructions are added to some prompts.
No cleaning has been done on the dataset… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Gemini-3.1-Pro-SmallWiki.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/gemini-3.1-pro-hard-high-reasoning.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gemini-3.1-pro-hard-high-reasoning.gemini-3.1-opus-4.6-reasoning-merged Merged from
https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning
https://huggingface.co/datasets/reedmayhew/gemini-3.1-pro-2048-reasoning-1100x
https://huggingface.co/datasets/crownelius/Opus-4.6-Reasoning-3300x
Gemini_3.1_202_Task_AI_Exposure_Scores
Gemini 3.1 2026 Task AI Exposure Scores
Dataset Summary
This dataset contains task-level AI exposure labels for O*NET task statements. Each task is classified into one of four categories, E0, E1, E2, or E3, using an updated 2026 Agentic AI Exposure Rubric and a Gemini 3.1 Pro classification pipeline. The labels are designed to capture whether a task can be accelerated by a frontier agentic AI system directly, whether it would require deeper software integration, or… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/Gemini_3.1_202_Task_AI_Exposure_Scores.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/REXX-NEW/gemini-3.1-pro-hard-high-reasoning.gemini-3.1-opus-4.6-reasoning-merged_v2gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/PhantomG27249/gemini-3.1-pro-hard-high-reasoning.gemini-3.1-pro-2048-reasoning-1100xgemini-3.1-opus-4.6-reasoning-merged Merged from
https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning
https://huggingface.co/datasets/reedmayhew/gemini-3.1-pro-2048-reasoning-1100x
https://huggingface.co/datasets/crownelius/Opus-4.6-Reasoning-3300x
gemini-3.1-pro-2048-reasoning-1100xgemini-3.1-pro-2048-reasoning-1100xgemini-3.1-opus-4.6-reasoning-merged Merged from
https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning
https://huggingface.co/datasets/reedmayhew/gemini-3.1-pro-2048-reasoning-1100x
https://huggingface.co/datasets/crownelius/Opus-4.6-Reasoning-3300x
