datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.auditbench-activations-jlens-NLA
AuditBench activations, J-lens readouts and NLA verbalizations
Every token of every AuditBench prompt and every model response, from
meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell.
Responses were regenerated greedily and run to the model's own stopping point rather
than truncated at a fixed length, and the activations, readouts and verbalizations
cover the prompt as well as the response.
84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.voice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.tulu3
ActiveUltraFeedback — Tulu 3
This is a preference dataset of 272k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026).
The prompts are from Tulu 3 8B Preference Mixture (Lambert et al., 2025). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/tulu3.gpio-llm-rpi5-actions
GPIO-LLM: Raspberry Pi 5 GPIO request-to-action dataset
Requests to a Raspberry Pi 5 in plain English, paired with the structured, validated GPIO action a
small on-device model should produce: a hardware operation, a clarifying question when the pin or device
is unknown, or a refusal when the request is invalid or unsafe. It was built to train a ~20M-parameter
English model that runs offline on the Pi.
Safety. Model output must never drive hardware directly. Every action is… See the full description on the dataset page: https://huggingface.co/datasets/AwaleSagar/gpio-llm-rpi5-actions.jma-gsi-disaster-action-corpus
JMA-GSI Disaster Action Corpus
A grounded, multilingual disaster-response dataset built from official Japanese government open data (JMA alert XML + JMA multilingual glossary + JMA forecast-area GIS + GSI designated evacuation shelters). Structured hazard alerts are transformed into easy-Japanese and multilingual (ja / easy-ja / en / vi / id / ne / my) action guidance, linked to hazard-compatible evacuation shelters, with full source traceability.
License (derived dataset): CC BY… See the full description on the dataset page: https://huggingface.co/datasets/edomaru/jma-gsi-disaster-action-corpus.skywork
ActiveUltraFeedback — Skywork
This is a preference dataset of 80k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026).
The prompts are from Skywork Reward Preference 80k v0.2 (Liu et al., 2024). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/skywork.ultrafeedback
ActiveUltraFeedback — UltraFeedback
This is a preference dataset of 60k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026).
The prompts are from the version of UltraFeedback (Cui et al., 2023) that was released by AllenAI (at allenai/ultrafeedback_binarized_cleaned). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/ultrafeedback.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.egocentric-activity-sample
Egocentric Activity Sample Dataset
A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks.
Dataset Summary
Metric
Value
Video clips
19
Total duration
~9.5 minutes
Resolution
960x540 (540p)
FPS
30
Narrations
99
NLQ queries
57
Moment annotations
19
FHO actions
57
Total size
~54 MB
Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.JL-ActionBoundary-1K-v0.1.0
JL-ActionBoundary-1K
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K is a 1,000-record English dataset for training and evaluating a narrow but important coding-agent behavior:
Before changing code, should the agent act, inspect the repository, ask the user, or defer because authority is missing?
The dataset is part of the JumpLander research direction on coding-agent behavior, repository intelligence, tool use, and controllable… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v0.1.0.ACT-2.5B
ACT-2.5B
Preview release: 183 sessions. This repository holds a preview of the dataset
described in a paper currently under double-blind review, so that the format and the
content can be inspected. It was drawn as a 200-session sample, of which 17 were withheld
during pre-release screening: 15 after a review of embedded images and 2 after an
identifier scan. The full subset of 26,999 sessions is released on
acceptance, once the redaction audit is complete. The paper's appendix… See the full description on the dataset page: https://huggingface.co/datasets/submit2331/ACT-2.5B.computer-use-large-actions
computer-use-large-actions
9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video).
Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0.
Split by software
category
examples
vscode
2,500
autocad
2,500
blender
1,000
excel
1,000
photoshop
1,000
salesforce
1,000
VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.when-agents-act
Dataset Card for "When Agents Act"
Dataset Summary
This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute).
Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.JudgeBias-DPO-RefFree-LOSO-action
JudgeBias-DPO-RefFree-LOSO-action
A Leave-One-Swap-Type-Out (LOSO) variant of JudgeBias-DPO-RefFree for evaluating out-of-distribution generalization of DPO-trained LLM judges.
Held-out Swap Type
Field
Value
Swap type
Action
Axis
τ⁻ (error)
Removed dataset
action_antonym_100pct
Description
action antonym substitution
All DPO pairs derived from action_antonym_100pct have been removed from both training and validation splits. The model trained… See the full description on the dataset page: https://huggingface.co/datasets/iknow-lab/JudgeBias-DPO-RefFree-LOSO-action.US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/actixon/US-PD-Books.vision-language-action-papers
Vision-Language-Action (VLA) & Robot Learning Papers — FineSet
A research-paper dataset on Vision-Language-Action (VLA) & Robot Learning Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Vision-Language-Action (VLA) & Robot Learning Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/vision-language-action-papers.code-action-sociale-familles
Code de l'action sociale et des familles, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-action-sociale-familles.Active_Listening_Content_1
Active-Listening-Content-1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Active_Listening_Content_1.
