datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jma-gsi-disaster-action-corpus
JMA-GSI Disaster Action Corpus
A grounded, multilingual disaster-response dataset built from official Japanese government open data (JMA alert XML + JMA multilingual glossary + JMA forecast-area GIS + GSI designated evacuation shelters). Structured hazard alerts are transformed into easy-Japanese and multilingual (ja / easy-ja / en / vi / id / ne / my) action guidance, linked to hazard-compatible evacuation shelters, with full source traceability.
License (derived dataset): CC BY… See the full description on the dataset page: https://huggingface.co/datasets/edomaru/jma-gsi-disaster-action-corpus.gpio-llm-rpi5-actions
GPIO-LLM: Raspberry Pi 5 GPIO request-to-action dataset
Requests to a Raspberry Pi 5 in plain English, paired with the structured, validated GPIO action a
small on-device model should produce: a hardware operation, a clarifying question when the pin or device
is unknown, or a refusal when the request is invalid or unsafe. It was built to train a ~20M-parameter
English model that runs offline on the Pi.
Safety. Model output must never drive hardware directly. Every action is… See the full description on the dataset page: https://huggingface.co/datasets/AwaleSagar/gpio-llm-rpi5-actions.task131_scan_long_text_generation_action_command_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.1.0
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
72 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.task127_scan_long_text_generation_action_command_all
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task127_scan_long_text_generation_action_command_all
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task127_scan_long_text_generation_action_command_all.code-as-action
Code-as-Action
Synthetic multi-turn trajectories for training language-model agents to solve problems by writing and executing code, then conditioning subsequent steps on execution feedback.
This release is the training corpus used for the LoRA adapter True2456/Qwen3.6-35B-A3B-Code-as-Action-LoRA.
Intent
The dataset operationalizes the CodeAct design pattern (Wang et al., ICML 2024): treat executable code as a unified action space for LLM agents, rather than… See the full description on the dataset page: https://huggingface.co/datasets/True2456/code-as-action.jam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.6.0 — a correction release. It withdraws 58 records whose source arrangements could not be licence-cleared and changes no remaining record. See Version 0.6.0 correction.
Records built: 2026-07-11 (0.5.0 cut; unchanged) Package built: 2026-09-25
DOI: 10.5281/zenodo.22961580 (this version; concept DOI 10.5281/zenodo.22961579). Earlier versions: 0.5.0 10.5281/zenodo.21313954 and 0.4.3 10.5281/zenodo.20279919. Both contain… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.task128_scan_structured_text_generation_command_action_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.zarn-meeting-to-actions
Zarn Meeting to Actions
Dataset Description
Meeting transcripts and notes mapped to summaries, decisions, owners, deadlines, and follow-up drafts.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem Need… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-meeting-to-actions.task129_scan_long_text_generation_action_command_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task129_scan_long_text_generation_action_command_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task129_scan_long_text_generation_action_command_short.JL-ActionBoundary-1K-v0.1.0
JL-ActionBoundary-1K
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K is a 1,000-record English dataset for training and evaluating a narrow but important coding-agent behavior:
Before changing code, should the agent act, inspect the repository, ask the user, or defer because authority is missing?
The dataset is part of the JumpLander research direction on coding-agent behavior, repository intelligence, tool use, and controllable… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v0.1.0.task130_scan_structured_text_generation_command_action_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.Fine_Grained_Fandom_Benchmark_Action_Sequences
Codified Decision Tree (CDT) Action Sequences
This dataset contains scene-action pairs derived from storylines, used to train and evaluate role-playing (RP) agents using the Codified Decision Trees (CDT) framework.
Paper: Deriving Character Logic from Storyline as Codified Decision Trees
Repository: https://github.com/KomeijiForce/Codified_Decision_Tree
Introduction
Role-playing (RP) agents rely on behavioral profiles to act consistently across diverse narrative… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Fine_Grained_Fandom_Benchmark_Action_Sequences.omnimcp_nextjs_server_actions_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_server_actions_teaser.mobile-actions-ita
Dataset Card: Mobile Actions (Italian Adaptation for Function Calling)
Overview
This dataset is an Italian adaptation of the original Google Mobile Actions dataset, designed to train lightweight models for on-device function calling. It preserves the original tool-calling schema in English while translating user interactions and contextual instructions into Italian.
The goal is to enable models to map natural language instructions in Italian to structured function calls… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/mobile-actions-ita.agentic-foresight-actions-2k
Agentic Foresight: 2K Multi-Step JSON Action & Rollback Dataset
Dataset Description
This dataset contains 2,000 highly structured, synthetically generated input/output pairs explicitly designed to train Large Language Models in Agentic Foresight, Multi-Step Orchestration, and Sequential Task Automation.
Unlike standard tool-calling datasets that map a single prompt to a single API call, this dataset forces the model to act as a macro-orchestrator. It translates… See the full description on the dataset page: https://huggingface.co/datasets/Qapdex/agentic-foresight-actions-2k.browser-agent-phase1-sft-action-only
Browser Agent Phase 1 SFT Action-Only
What this is
Action-only step-level chat SFT data for browser-agent training.
Each example teaches the model to predict the next BrowserGym action from:
the original generation-time system prompt used for data collection
task goal and URL
short recent history
current observation text and diagnostics
Assistant targets contain only the next action.
Why this format
This is the primary training format for small-model SFT… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-action-only.computer-use-large-actions
computer-use-large-actions
9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video).
Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0.
Split by software
category
examples
vscode
2,500
autocad
2,500
blender
1,000
excel
1,000
photoshop
1,000
salesforce
1,000
VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.JudgeBias-DPO-RefFree-LOSO-action
JudgeBias-DPO-RefFree-LOSO-action
A Leave-One-Swap-Type-Out (LOSO) variant of JudgeBias-DPO-RefFree for evaluating out-of-distribution generalization of DPO-trained LLM judges.
Held-out Swap Type
Field
Value
Swap type
Action
Axis
τ⁻ (error)
Removed dataset
action_antonym_100pct
Description
action antonym substitution
All DPO pairs derived from action_antonym_100pct have been removed from both training and validation splits. The model trained… See the full description on the dataset page: https://huggingface.co/datasets/iknow-lab/JudgeBias-DPO-RefFree-LOSO-action.authority-to-action
Authority-to-Action Evaluation
Research question. When relevant context is present, does a tool-using
language-model system distinguish evidence from permission to act?
This dataset contains the 100 synthetic cases and the transcript-free,
600-row results ledger behind the LatentAtlas Authority-to-Action study.
Paper (preprint): doi:10.5281/zenodo.21957491
Code, Inspect evaluation and deterministic verifiers:
github.com/latentatlas/latentatlas-evidence-evals
BENCHMARK DATA… See the full description on the dataset page: https://huggingface.co/datasets/hbuldurgan/authority-to-action.Bandori_Conversational_Benchmark_Action_Sequences
Codified Decision Tree (CDT) Dataset
Paper | GitHub
Codified Decision Trees (CDT) is a framework that induces executable and interpretable behavioral profiles for role-playing (RP) agents from narrative data. This dataset contains scene-action pairs derived from diverse storylines used to construct and validate these behavioral representations.
Introduction
Role-playing agents often rely on unstructured profiles that lead to brittle behavior. CDT represents behavioral… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Bandori_Conversational_Benchmark_Action_Sequences.browser-agent-phase1-sft-reasoning-action
Browser Agent Phase 1 SFT Reasoning+Action
What this is
Reasoning-plus-action step-level chat SFT data for browser-agent training.
Each example uses the original generation-time system prompt, then appends a short instruction to reason first and output the final action.
Assistant targets contain:
one <think>...</think> block
then one BrowserGym action
Why this format
This is an experimental variant for comparing whether explicit reasoning supervision helps or… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-reasoning-action.counterfactual-action-invariants-v0.1
What this dataset tests
Leaders demand causality.
Reality gives entanglement.
You must keep invariants.
Why it exists
Models often answer a forced question.
They pick one cause.
They fake proof.
This set checks whether you
resist false certainty
name confounders
propose a valid counterfactual method
turn pressure into a decision gate
Data format
Each row contains
scenario_context
user_message
counterfactual_pressure
constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/counterfactual-action-invariants-v0.1.language-action-RVQ-CoT-humanML
A chain-of-thought dataset for human movement generation using a RVQ.
Generated from: https://huggingface.co/Wojtekb30/HumanML3D-500ms-FPP-descriptions-CoTs-1
With use of: https://huggingface.co/Wojtekb30/Motion-RVQ-263d-reconstructor-humanML
This dataset can be used to train LLMs into generating chains of thought which contain RVQ movement tokens, allowing generation of human movements.
arabic-mobile-actions
Arabic Mobile Actions Dataset
Arabic Function Calling dataset reformatted for FunctionGemma fine-tuning.
This dataset converts the Arabic Function Calling Dataset to the Google Mobile Actions format, enabling fine-tuning of FunctionGemma for Arabic on-device function calling.
Dataset Statistics
Metric
Value
Total samples
45,729
Train samples
36,583 (80%)
Eval samples
9,146 (20%)
Positive (with tool call)
41,175
Negative (no tool call)
4,554… See the full description on the dataset page: https://huggingface.co/datasets/Sa74ll/arabic-mobile-actions.trust-safety-action-routing
Trust and Safety Action Routing
This dataset evaluates moderation behavior for UGC, marketplaces, direct
messages, and moderation queues. It focuses on action routing: rewrite hostile
or policy-violating content, redact PII, refuse coordinated abuse, escalate
credible threats, and test shadow-mode rollout.
The dataset reflects a practical trust and safety requirement: platforms often
need an action and a reason code, not just a harmful/not-harmful label.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/abliterationaiorg/trust-safety-action-routing.embodied-action-outcome-coherence-v0.1Embodied Action–Outcome Coherence v0.1
What this tests
Whether an embodied agent updates world state from observed outcomes
Whether it avoids claiming success when the outcome says failure
Failure modes
outcome_ignoredResponse does not reflect the true post-action state
false_successResponse claims success despite an observed failure
causal_update_okResponse states the correct post-action state without contradiction
How it works
world_facts_t0 is the initial state
action_taken is what the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-action-outcome-coherence-v0.1.tool-reasoning-sft-TOOLS-mobile-actions-data-cleaned-rectified
Mobile Actions — Cleaned & Rectified
8.7K on-device function calling conversations converted into a strict reasoning + tool-call format. Covers 7 Android mobile actions including calendar events, emails, contacts, maps, flashlight, and Wi-Fi settings.
Format
Each row contains a structured conversation with explicit reasoning traces and validated tool calls.
Message Roles
Role
Content
system
Tool-use protocol + cleaned JSON tool schemas +… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-mobile-actions-data-cleaned-rectified.
