datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mobile-actions
Mobile Actions: A Dataset for On-Device Function Calling
The dataset contains conversational traces designed to train lightweight models (such as FunctionGemma 270M) to translate natural language instructions into executable function calls for Android OS system tools.
Dataset Format
The dataset is provided in JSONL format. Each line represents a data sample. The
dataset is pre-split into training and evaluation sets. This distinction is
denoted by the metadata field… See the full description on the dataset page: https://huggingface.co/datasets/google/mobile-actions.human_assisted_action_preference_optimizationgfmc_hyworld1.5_processed_160latents_16fps_actionaction-atlas-oft-activationsaction-roleplay-data
Action Roleplay Data
Data package for the Action SA-MP Android client.
The client connects to 92.119.165.177:5636. The files/ directory contains the extracted game data, cache.zip is the archive consumed by the initial installer, files.json is the file-by-file manifest, and client_config.json contains the public endpoints. Runtime logs were excluded from the distributable package.
The APK included here is a debug build for testing and is signed with a debug key.
github-actionsjam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.2026-08-28-post-action-retrospection-716-coherent
Post-action retrospection 716 -- coherent rewrite (arm 1 of the PAR coherence experiment)
field
value
experiment
The exact 716 five-turn PAR rows that trained LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch (mixture 2026-08-26-table2-9284-par716-train @ 42c8a74), with ONLY the trained turn (turn 4: private reasoning + reply) rewritten by Sonnet 5 so the reasoning ENDS on a first-person decision (what it won't do, per… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-post-action-retrospection-716-coherent.action15s-media-20260910Media and bilingual annotations for the companion gameplay action review.
Use train_15s.jsonl as the current accepted selection: each row contains its relative video path, checksum, source interval, and English/Chinese timed action labels. Historical media files from earlier progress snapshots may remain in the repository; only the manifest defines the current batch. summary.json reports coverage and pending visual screening separately. The default dataset configuration reads only the training… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action15s-media-20260910.jam-actions-v1-probe
jam-actions-v1-probe
Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only ·
Companion to: jam-actions-v1
Why it exists
An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes.
Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that
corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.jam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.5.1 — a documentation-only revision of the 0.5.0 record cut. No record, split or eval artifact changed; the card gained the fine-tuning evaluation banner and the "What's in a record" walkthrough, which had been added on Hugging Face and lived nowhere else.
Records built: 2026-07-11 Source tag: jam-actions-v0-0.5.0-cut-2026-07-11 (record-content correction release — Bach BWV 846 errata 001 + 002; see RELEASE_NOTES.md… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.0.2
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.code-as-action
Code-as-Action
Synthetic multi-turn trajectories for training language-model agents to solve problems by writing and executing code, then conditioning subsequent steps on execution feedback.
This release is the training corpus used for the LoRA adapter True2456/Qwen3.6-35B-A3B-Code-as-Action-LoRA.
Intent
The dataset operationalizes the CodeAct design pattern (Wang et al., ICML 2024): treat executable code as a unified action space for LLM agents, rather than… See the full description on the dataset page: https://huggingface.co/datasets/True2456/code-as-action.cs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.zarn-meeting-to-actions
Zarn Meeting to Actions
Dataset Description
Meeting transcripts and notes mapped to summaries, decisions, owners, deadlines, and follow-up drafts.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem Need… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-meeting-to-actions.ActionFlow-CodeBank-v1JL-ActionBoundary-1K-v0.1.0
JL-ActionBoundary-1K
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K is a 1,000-record English dataset for training and evaluating a narrow but important coding-agent behavior:
Before changing code, should the agent act, inspect the repository, ask the user, or defer because authority is missing?
The dataset is part of the JumpLander research direction on coding-agent behavior, repository intelligence, tool use, and controllable… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v0.1.0.Fence-Climbing-Action-Recognition-Dataset
Fence Climbing Action Recognition Dataset
The current security industry faces challenges from people climbing over walls, fences, and other security hazards. Traditional surveillance methods often cannot timely and effectively recognize these abnormal behaviors. Existing solutions are insufficient in the accuracy and real-time detection of actions, resulting in the inability to quickly respond to potential dangers. This dataset aims to support the training of action recognition… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Fence-Climbing-Action-Recognition-Dataset.2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.ovdsgg-action-genome-split
OvDSGG Action Genome Open-Vocabulary Split
This dataset repository contains the open-vocabulary category split used by OvDSGG for Action Genome experiments.
It does not redistribute Action Genome videos, frames, or full annotations. Users must obtain and process Action Genome separately, then use this split metadata to reproduce the OvDSGG open-vocabulary training/evaluation protocol.
Paper: https://huggingface.co/papers/2608.14835Code: https://github.com/jhelsby/OvDSGGModel… See the full description on the dataset page: https://huggingface.co/datasets/jhelsby/ovdsgg-action-genome-split.agentic-foresight-actions-2k
Agentic Foresight: 2K Multi-Step JSON Action & Rollback Dataset
Dataset Description
This dataset contains 2,000 highly structured, synthetically generated input/output pairs explicitly designed to train Large Language Models in Agentic Foresight, Multi-Step Orchestration, and Sequential Task Automation.
Unlike standard tool-calling datasets that map a single prompt to a single API call, this dataset forces the model to act as a macro-orchestrator. It translates… See the full description on the dataset page: https://huggingface.co/datasets/Qapdex/agentic-foresight-actions-2k.vcore-actions-benchmark
Authority envelopes of AI agents in GitHub workflows: benchmark
Links
Blog post: https://paulinebourigault.github.io/blog/2026/what-the-agent-may-do/
Code (vcore): https://github.com/certior/vcore
Companion dataset (corpus): https://huggingface.co/datasets/paulibo/vcore-workflow-envelopes
3,229 runs of GitHub Actions workflows that use
anthropics/claude-code-action, pinned
at v1.0.225 (bf38e86). Each run pairs what the action itself did with what an
authority… See the full description on the dataset page: https://huggingface.co/datasets/paulibo/vcore-actions-benchmark.mobile-actions-ita
Dataset Card: Mobile Actions (Italian Adaptation for Function Calling)
Overview
This dataset is an Italian adaptation of the original Google Mobile Actions dataset, designed to train lightweight models for on-device function calling. It preserves the original tool-calling schema in English while translating user interactions and contextual instructions into Italian.
The goal is to enable models to map natural language instructions in Italian to structured function calls… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/mobile-actions-ita.tb21-eval-qwen35-action-only-40k-c164-max32k-timeout8x
qwen35-action-only-40k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-40k-tacc through the served
model ID qwen35-action-only-40k with Terminus-2.
Noncanonical run: timeout_multiplier=8 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Recorded trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 111
Exception… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-40k-c164-max32k-timeout8x.actionnet-grasp-cosmoscomputer-use-large-actions
computer-use-large-actions
9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video).
Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0.
Split by software
category
examples
vscode
2,500
autocad
2,500
blender
1,000
excel
1,000
photoshop
1,000
salesforce
1,000
VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x
qwen35-action-only-20k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-20k-tacc through the served
model ID qwen35-action-only-20k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; repair concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Recorded trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 250… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-20k-infra-repaired-c164-max32k-timeout2x.catrace-teen-stress-coping-actions
CatRace v0.2 development dataset
CatRace v0.2 contains 5,100 feasibility records and 3,600 pairwise-ranking records.
Quick start
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python scripts/train_feasibility_baseline.py
python scripts/train_barrier_baseline.py
Structure
data/: train, validation, test, and 200-record human-audit files in JSONL and CSV
pairs/: pairwise ranking data
scripts/: fast baseline models… See the full description on the dataset page: https://huggingface.co/datasets/nyasaluharuka/catrace-teen-stress-coping-actions.tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x
qwen35-action-only-10k — Terminal-Bench 2.1 (timeout multiplier 2x)
Terminal-Bench 2.1 evaluation protocol variant (timeout multiplier 2x) of violetxi/qwen35-4b-offline-echo-action-only-10k-tacc through the served
model ID qwen35-action-only-10k with Terminus-2.
Result
Evaluation trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 219
Agent timeouts / context-length events / output-cap events:
217 / 0 /
0
Mean reward / Pass@1: 0.105618
Pass@5:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x.
