datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Horizon-Execution
Long Horizon Execution
This project contains the dataset accompanying the paper "The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs"
Abstract
Does continued scaling of large language models (LLMs) yield diminishing returns? Real-world value often stems from the length of task an agent can complete. We start this work by observing the simple but counterintuitive fact that marginal gains in single-step accuracy can compound into exponential… See the full description on the dataset page: https://huggingface.co/datasets/arvindh75/Long-Horizon-Execution.execution-verified-codework
Execution-Verified CodeWork
Only code that passes the tests ships
Sandbox-executed · ≥6 unit tests · implement / repair / harden · instance-deduplicated
One-sentence pitch
Training traces for writing, fixing, and hardening Python functions — every kept solution was actually run against unit tests and passed.
What you get
Field
Role
kind
implement · repair · harden
problem
Clear developer task
reasoning
Numbered… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/execution-verified-codework.code-execution-trace-training-pool
Code execution trace training pool
Public Python code paired with one concrete call and the value that call returns. Every value in
this pool was computed by running the code, not copied from a label. The data is laid out twice,
and either layer may be used.
pool/
Every source rewritten into one shape, 2174322 rows over 11 gzipped parts, one JSON object per
line, with these fields.
Field
What it holds
id
a row identifier unique within this pool
code… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/code-execution-trace-training-pool.python-execution-prediction-training-pool
Python execution prediction training pool
Short Python functions, a concrete call of each one, and the value that call really returns, from
five public sources read at the pinned revisions named below and laid out twice. Train on either
layer or on both.
pool.jsonl
Every source rewritten into one shape, 38154 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
code
the Python source that… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-execution-prediction-training-pool.code-procedures-civiles-execution
Code des procédures civiles d'exécution, non-instruct (2025-03-10)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-procedures-civiles-execution.python-mental-execution-traces
Python Mental Execution Traces
A 12,000-row prompt/completion dataset for evaluating and training language models to mentally execute self-contained Python 3 snippets without running them. Completions provide the expected standard output together with a concise variable trace or explanation.
Dataset structure
The JSONL file contains two text fields:
prompt: a Python mental-execution problem.
completion: the expected stdout and concise reasoning or variable trace.… See the full description on the dataset page: https://huggingface.co/datasets/ILoveBuns/python-mental-execution-traces.terminal-command-execution-sft
Terminal Command Execution SFT
A merged conversational SFT dataset for training careful terminal command assistants across POSIX shells, Linux, macOS, WSL, Termux, Windows Command Prompt, PowerShell, Nushell, Docker, Git, package managers, process inspection, system inspection, and scripting/control-flow tasks.
Format
Each row follows a TRL/Unsloth-compatible conversational format:
{
"messages": [
{
"role": "system",
"content": "You are a careful… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/terminal-command-execution-sft.task-execution-quality
TASK_EXECUTION_QUALITY
A preference dataset for TASK_EXECUTION_QUALITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/task-execution-quality.optimal-execution-route-prefs
OPTIMAL_EXECUTION_ROUTE
A preference dataset for OPTIMAL_EXECUTION_ROUTE, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally query)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/optimal-execution-route-prefs.Deterministic-Execution-Data-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Deterministic-Execution-Data-Layer.Execution-Bound-Artifact-Reconstruction-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Execution-Bound-Artifact-Reconstruction-Layer.execution-verified-agent-trajectories
Execution-Verified Agent Trajectories — Format & Method
This repository documents a method and data format for building supervised fine-tuning sets from agent
trajectories that are verified by running the code, not by asking a model whether the answer looks right.
This is a specification plus synthetic examples, not a corpus. The trajectories that trained
Luthor 8B were generated against a private repository and cannot be
released. Everything needed to rebuild an equivalent set… See the full description on the dataset page: https://huggingface.co/datasets/IAMIbrahim/execution-verified-agent-trajectories.Execution-Bound-Artifact-Reconstruction-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Execution-Bound-Artifact-Reconstruction-Layer.Deterministic-Execution-Data-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Deterministic-Execution-Data-Layer.Machine-Checkable-Blockchain-Execution-Specification
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Machine-Checkable-Blockchain-Execution-Specification.Machine-Checkable-Blockchain-Execution-Specification
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Machine-Checkable-Blockchain-Execution-Specification.
