datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
action-atlas-oft-activationsgithub-actionsaction15s-media-20260910Media and bilingual annotations for the companion gameplay action review.
Use train_15s.jsonl as the current accepted selection: each row contains its relative video path, checksum, source interval, and English/Chinese timed action labels. Historical media files from earlier progress snapshots may remain in the repository; only the manifest defines the current batch. summary.json reports coverage and pending visual screening separately. The default dataset configuration reads only the training… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action15s-media-20260910.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.cs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.JL-ActionBoundary-1K-v0.1.0
JL-ActionBoundary-1K
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K is a 1,000-record English dataset for training and evaluating a narrow but important coding-agent behavior:
Before changing code, should the agent act, inspect the repository, ask the user, or defer because authority is missing?
The dataset is part of the JumpLander research direction on coding-agent behavior, repository intelligence, tool use, and controllable… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v0.1.0.2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.actionnet-grasp-cosmostb21-eval-qwen35-action-only-40k-c164-max32k-timeout8x
qwen35-action-only-40k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-40k-tacc through the served
model ID qwen35-action-only-40k with Terminus-2.
Noncanonical run: timeout_multiplier=8 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Recorded trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 111
Exception… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-40k-c164-max32k-timeout8x.computer-use-large-actions
computer-use-large-actions
9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video).
Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0.
Split by software
category
examples
vscode
2,500
autocad
2,500
blender
1,000
excel
1,000
photoshop
1,000
salesforce
1,000
VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x
qwen35-action-only-20k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-20k-tacc through the served
model ID qwen35-action-only-20k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; repair concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Recorded trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 250… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x.tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x
qwen35-action-only-10k — Terminal-Bench 2.1 (timeout multiplier 2x)
Terminal-Bench 2.1 evaluation protocol variant (timeout multiplier 2x) of violetxi/qwen35-4b-offline-echo-action-only-10k-tacc through the served
model ID qwen35-action-only-10k with Terminus-2.
Result
Evaluation trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 219
Agent timeouts / context-length events / output-cap events:
217 / 0 /
0
Mean reward / Pass@1: 0.105618
Pass@5:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x.sonar-municipal-pl-actions
Sonar Municipal — PL Actions Corpus
The largest publicly released Portuguese legal text rewriting dataset: 241,111 pairs mapping the original ementa (summary) of a Brazilian municipal Projeto de Lei (PL) to its action-form textualization (ação). Action form is a direct, imperative rewrite that strips juridical boilerplate, normalizes typography, and surfaces the underlying intervention rather than the legal instrument.
Companion to: ICMC-USP undergraduate thesis "Mineração… See the full description on the dataset page: https://huggingface.co/datasets/thiagoambiel/sonar-municipal-pl-actions.action2code
RoboCasa Pretrain65 Action2Code: Raw Extracted and Compressed
Self-contained local dataset for the 65 RoboCasa pretrain atomic tasks. Each task has two rows:
raw_extracted for the pre-compression open-loop-joint code and compressed for the final source-compressed code.
Videos are copied into the dataset under videos/<dataset_type>/ and use the latest whole-body-IK validation rerun.
Common Fields
dataset_type: raw_extracted or compressed.
pretrain_task_id… See the full description on the dataset page: https://huggingface.co/datasets/stovecat/action2code.action-evidence-vla-phase-state-cachelatent-action-can-mh-image-v15
joon-stack/latent-action-can-mh-image-v15
Raw HDF5 artifact mirror for latent_action training.
This repository is not a native load_dataset(...) dataset.
It stores robomimic-style HDF5 files for download and local path overrides.
Contents
Files: 1
Total bytes: 12452033552
Usage
Download the needed HDF5 file locally and pass it to Hydra:
python run.py data.paths=[/absolute/path/to/can/ph/image_v15.hdf5]
Files
can/mh/image_v15.hdf5… See the full description on the dataset page: https://huggingface.co/datasets/quiet-storm/latent-action-can-mh-image-v15.tb21-eval-qwen35-4b-action-only-20k-thinking-timeout2x-gcp4-infra-interrupted
qwen35-action-only-20k — Terminal-Bench 2.1
Nonstandard Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-20k-tacc@bd9c914f4774b2f217cecdf1ed18a7d2f3e0a623 through the served
model ID qwen35-action-only-20k with Terminus-2.
Result
Recorded trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 408
Exception counts: {"AgentTimeoutError": 41, "InternalServerError": 367}
Agent timeouts / context-length events / output-cap… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-action-only-20k-thinking-timeout2x-gcp4-infra-interrupted.vision-language-action-papers
Vision-Language-Action (VLA) & Robot Learning Papers — FineSet
A research-paper dataset on Vision-Language-Action (VLA) & Robot Learning Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Vision-Language-Action (VLA) & Robot Learning Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/vision-language-action-papers.tb21-eval-qwen35-action-only-40k-c164-max32k-timeout2x
qwen35-action-only-40k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-40k-tacc through the served
model ID qwen35-action-only-40k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Recorded trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 211
Exception… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-40k-c164-max32k-timeout2x.han-humanoid-adaptive-action-feedback-v1
Humanoid Adaptive Action Feedback
Overview
This dataset contains structured feedback
on executed humanoid actions and subsequent adjustments.
It supports reinforcement and adaptive behavior modeling.
Data Fields
action_id
executed_action
environment_state
human_feedback_score
adjustment_applied
post_adjustment_score
Intended Use
Reinforcement learning research
Continuous improvement systems
Adaptive humanoid frameworks
License
MIT
latent-action-can-ph-image-v15
joon-stack/latent-action-can-ph-image-v15
Raw HDF5 artifact mirror for latent_action training.
This repository is not a native load_dataset(...) dataset.
It stores robomimic-style HDF5 files for download and local path overrides.
Contents
Files: 1
Total bytes: 4610179960
Usage
Download the needed HDF5 file locally and pass it to Hydra:
python run.py data.paths=[/absolute/path/to/can/ph/image_v15.hdf5]
Files
can/ph/image_v15.hdf5… See the full description on the dataset page: https://huggingface.co/datasets/quiet-storm/latent-action-can-ph-image-v15.p2-robocasa-action-adapter-rolloutsus-science-policy-bills
US Science Policy Bills
A stance-annotated corpus of 8,191 public-health-related bills from the United States federal Congress and 41 state legislatures, collected and classified by SAFE Action (Science and Freedom for Everyone Action Fund), a 501(c)(4) social welfare organization building open-source democratic tools for science-based policy.
To our knowledge this is the first openly published, stance-annotated dataset of state-level public health legislation. Raw legislative… See the full description on the dataset page: https://huggingface.co/datasets/SAFE-Action/us-science-policy-bills.cs001cs002cs006orbit-wars-value-v2anchor-actionscs007cs003cs005
