datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minecraft-text-action-datasetaction-atlas-groot-activationsBreakfast-Actions
🍳 Breakfast Actions Dataset (HF + WebDataset Ready)
This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows.It provides:
4 evaluation splits (s1, s2, s3, s4)
JSONL metadata describing each video, participant, camera, and frame-level action segments
Raw AVI videos stored directly on HuggingFace
Optional WebDataset shards for streaming training
📁 Folder Layout
Breakfast-Actions/
│
├──… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/Breakfast-Actions.action100m-preview
Action100M: A Large-scale Video Action Dataset
Paper | GitHub
Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling.
Load Action100M Annotations
Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.mobile-actions
Mobile Actions: A Dataset for On-Device Function Calling
The dataset contains conversational traces designed to train lightweight models (such as FunctionGemma 270M) to translate natural language instructions into executable function calls for Android OS system tools.
Dataset Format
The dataset is provided in JSONL format. Each line represents a data sample. The
dataset is pre-split into training and evaluation sets. This distinction is
denoted by the metadata field… See the full description on the dataset page: https://huggingface.co/datasets/google/mobile-actions.actionnet-subset100-gtdepth
ActionNet subset100 with ground-truth depth
This is a 100-episode subset redistribution of a third-party dataset, plus our derived
artifacts. Read the attribution before using it.
Attribution and license
Upstream dataset
FourierIntelligence/ActionNet (Fourier Intelligence)
What is redistributed
100 episodes out of 30,121, byte-identical to the upstream tars: rgb.mp4 (1280x800 fisheye), depth.mkv (lossless 16-bit), timestamps.json, <ULID>.hdf5… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/actionnet-subset100-gtdepth.human_assisted_action_preference_optimizationgfmc_hyworld1.5_processed_160latents_16fps_actionaction-atlas-oft-activationsminecraft-motion-action-datasetrobot-action-prediction-dataset
Robotic Action Prediction Dataset
Dataset Description
This dataset contains triplets of (current observation, action instruction, future observation) for training models to predict future frames of robotic actions.
Dataset Structure
Data Fields
current_frame: Input image (RGB) of the current observation
instruction: Textual description of the action to perform
future_frame: Target image (RGB) showing the expected outcome 50 frames later… See the full description on the dataset page: https://huggingface.co/datasets/bryandts/robot-action-prediction-dataset.action-roleplay-data
Action Roleplay Data
Data package for the Action SA-MP Android client.
The client connects to 92.119.165.177:5636. The files/ directory contains the extracted game data, cache.zip is the archive consumed by the initial installer, files.json is the file-by-file manifest, and client_config.json contains the public endpoints. Runtime logs were excluded from the distributable package.
The APK included here is a debug build for testing and is signed with a debug key.
github-actionsdrive-actionActionEQA
ActionEQA: Action Interface for Embodied Question Answering
Tianwei Bao1* · Qineng Wang1* · Kangrui Wang1 · Mingkai Deng2 · Guangyi Liu5 · Jiayuan Mao3
Larry Birnbaum1 · Zhiting Hu4 · Eric P. Xing2,5 · Zhaoran Wang1 · Manling Li1
1 Northwestern University 2 Carnegie Mellon University 3 UPenn
4 UC San Diego 5 MBZUAI
* Equal contribution
ActionEQA is the first action-centric Embodied Question Answering (EQA) benchmark designed to systematically evaluate… See the full description on the dataset page: https://huggingface.co/datasets/TianweiBao/ActionEQA.2026-08-27-odcv-post-action-retrospection-716-seed-2-eval
ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 2, 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-2-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-2-eval.minecraft-grounding-action-datasetactionWM_vfx_sample150
H3 VFX editing: 150 review examples
150 synthetic VFX event videos for colleague review, generated with MiniMax H3 Max
text-to-video and balanced prompt expansion. Each raw MP4 and its exact submitted
raw prompt TXT share a task ID and are stored together in this directory.
The text column in the dataset viewer contains the same raw prompt.
Category
First person
Third person
Total
Global environment
30
30
60
Local object
30
30
60
Character effects
0
30
30
Total… See the full description on the dataset page: https://huggingface.co/datasets/Geral-Yuan/actionWM_vfx_sample150.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.gambitflow-chess-actionvalues-v1-cleanjma-gsi-disaster-action-corpus
JMA-GSI Disaster Action Corpus
A grounded, multilingual disaster-response dataset built from official Japanese government open data (JMA alert XML + JMA multilingual glossary + JMA forecast-area GIS + GSI designated evacuation shelters). Structured hazard alerts are transformed into easy-Japanese and multilingual (ja / easy-ja / en / vi / id / ne / my) action guidance, linked to hazard-compatible evacuation shelters, with full source traceability.
License (derived dataset): CC BY… See the full description on the dataset page: https://huggingface.co/datasets/edomaru/jma-gsi-disaster-action-corpus.gpio-llm-rpi5-actions
GPIO-LLM: Raspberry Pi 5 GPIO request-to-action dataset
Requests to a Raspberry Pi 5 in plain English, paired with the structured, validated GPIO action a
small on-device model should produce: a hardware operation, a clarifying question when the pin or device
is unknown, or a refusal when the request is invalid or unsafe. It was built to train a ~20M-parameter
English model that runs offline on the Pi.
Safety. Model output must never drive hardware directly. Every action is… See the full description on the dataset page: https://huggingface.co/datasets/AwaleSagar/gpio-llm-rpi5-actions.corporate-actions
US Corporate Actions — dividends and splits
391 639 dividends from 3 327 filers · 5 619 splits from 3 814 filers ·
2005 to 2026
Built to close a specific hole. A filing states shares and earnings per share
as of the day it was made; every price series is adjusted for splits since.
Multiply one by the other and the answer is wrong by the split factor — on
Deckers that turned a 6.9% earnings yield into 41.7%, a P/E of 1.8.
The pipeline lives in recipe/ at the same revision as the… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/corporate-actions.2026-08-28-post-action-retrospection-716-coherent
Post-action retrospection 716 -- coherent rewrite (arm 1 of the PAR coherence experiment)
field
value
experiment
The exact 716 five-turn PAR rows that trained LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch (mixture 2026-08-26-table2-9284-par716-train @ 42c8a74), with ONLY the trained turn (turn 4: private reasoning + reply) rewritten by Sonnet 5 so the reasoning ENDS on a first-person decision (what it won't do, per… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-post-action-retrospection-716-coherent.action15s-media-20260910Media and bilingual annotations for the companion gameplay action review.
Use train_15s.jsonl as the current accepted selection: each row contains its relative video path, checksum, source interval, and English/Chinese timed action labels. Historical media files from earlier progress snapshots may remain in the repository; only the manifest defines the current batch. summary.json reports coverage and pending visual screening separately. The default dataset configuration reads only the training… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action15s-media-20260910.task131_scan_long_text_generation_action_command_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.tb2-eval-qwen3-8b-action-clean-6ep
Terminal-Bench 2.0 eval results - Qwen3-8B action-clean 6ep SFT
Compact Harbor eval results for violetxi/qwen3-8b-terminal-action-clean-6ep on
Terminal-Bench 2.0 using the terminus-2 agent harness.
Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward,
per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata.
Raw terminal recordings, panes, and completion logs are not included.… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-action-clean-6ep.jam-actions-v1-probe
jam-actions-v1-probe
Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only ·
Companion to: jam-actions-v1
Why it exists
An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes.
Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that
corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.1.0
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
72 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.pi0_stacking_action_chunks
