CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01actava /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.documenttext-generationn<1K61 likes6.2k downloads4mo agoHugging Face02PranavViswanath /auditbench-activations-jlens-NLA AuditBench activations, J-lens readouts and NLA verbalizations Every token of every AuditBench prompt and every model response, from meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell. Responses were regenerated greedily and run to the model's own stopping point rather than truncated at a fixed length, and the activations, readouts and verbalizations cover the prompt as well as the response. 84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.tabulartext-generation100M<n<1B0 likes3.2k downloads1mo agoHugging Face03laion /voice-acting-cutscene-prompts Cut-Scene Voice-Acting Prompts Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting models. Each prompt describes a single speaker across two sharply contrasting emotional moments separated by a CUT TO: transition, in a voice-acting stage-direction format (spoken lines in "quotes", performance notes in (parentheses)). Total prompts: 4,057,000 Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.tabulartext-generation1M<n<10M2 likes993 downloads12d agoHugging Face04ActiveUltraFeedback /tulu3 ActiveUltraFeedback — Tulu 3 This is a preference dataset of 272k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026). The prompts are from Tulu 3 8B Preference Mixture (Lambert et al., 2025). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/tulu3.tabulartext-generation1M<n<10M0 likes439 downloads4mo agoHugging Face05AwaleSagar /gpio-llm-rpi5-actions GPIO-LLM: Raspberry Pi 5 GPIO request-to-action dataset Requests to a Raspberry Pi 5 in plain English, paired with the structured, validated GPIO action a small on-device model should produce: a hardware operation, a clarifying question when the pin or device is unknown, or a refusal when the request is invalid or unsafe. It was built to train a ~20M-parameter English model that runs offline on the Pi. Safety. Model output must never drive hardware directly. Every action is… See the full description on the dataset page: https://huggingface.co/datasets/AwaleSagar/gpio-llm-rpi5-actions.tabulartext-generation1M<n<10M0 likes188 downloads4d agoHugging Face06edomaru /jma-gsi-disaster-action-corpus JMA-GSI Disaster Action Corpus A grounded, multilingual disaster-response dataset built from official Japanese government open data (JMA alert XML + JMA multilingual glossary + JMA forecast-area GIS + GSI designated evacuation shelters). Structured hazard alerts are transformed into easy-Japanese and multilingual (ja / easy-ja / en / vi / id / ne / my) action guidance, linked to hazard-compatible evacuation shelters, with full source traceability. License (derived dataset): CC BY… See the full description on the dataset page: https://huggingface.co/datasets/edomaru/jma-gsi-disaster-action-corpus.tabularquestion-answering100K<n<1M1 likes156 downloads5mo agoHugging Face07ActiveUltraFeedback /skywork ActiveUltraFeedback — Skywork This is a preference dataset of 80k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026). The prompts are from Skywork Reward Preference 80k v0.2 (Liu et al., 2024). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first generate candidate responses, then uses various active selection strategies… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/skywork.tabulartext-generation100K<n<1M0 likes135 downloads4mo agoHugging Face08ActiveUltraFeedback /ultrafeedback ActiveUltraFeedback — UltraFeedback This is a preference dataset of 60k samples generated for the paper ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning (Melikidze et al., 2026). The prompts are from the version of UltraFeedback (Cui et al., 2023) that was released by AllenAI (at allenai/ultrafeedback_binarized_cleaned). The response pairs were generated with the ActiveUltraFeedback pipeline, which calls a large pool of open-weight LLMs to first… See the full description on the dataset page: https://huggingface.co/datasets/ActiveUltraFeedback/ultrafeedback.tabulartext-generation100K<n<1M0 likes131 downloads4mo agoHugging Face09jumplander /JL-ActionBoundary-1K-v1.0.0 JL-ActionBoundary-1K v1.0.0 Counterfactual Ask–Inspect–Act–Defer supervision for coding agents JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code: ACT: the task is sufficiently specified for bounded repository work; INSPECT: missing information can be recovered from the repository; ASK: a material product decision belongs to the user; DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.tabulartext-classification1K<n<10K5 likes122 downloads2mo agoHugging Face10WhissleAI /egocentric-activity-sample Egocentric Activity Sample Dataset A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks. Dataset Summary Metric Value Video clips 19 Total duration ~9.5 minutes Resolution 960x540 (540p) FPS 30 Narrations 99 NLQ queries 57 Moment annotations 19 FHO actions 57 Total size ~54 MB Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.tabularvideo-classificationn<1K0 likes108 downloads5mo agoHugging Face11jumplander /JL-ActionBoundary-1K-v0.1.0 JL-ActionBoundary-1K Counterfactual Ask–Inspect–Act–Defer supervision for coding agents JL-ActionBoundary-1K is a 1,000-record English dataset for training and evaluating a narrow but important coding-agent behavior: Before changing code, should the agent act, inspect the repository, ask the user, or defer because authority is missing? The dataset is part of the JumpLander research direction on coding-agent behavior, repository intelligence, tool use, and controllable… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v0.1.0.tabulartext-classification1K<n<10K5 likes86 downloads2mo agoHugging Face12submit2331 /ACT-2.5B ACT-2.5B Preview release: 183 sessions. This repository holds a preview of the dataset described in a paper currently under double-blind review, so that the format and the content can be inspected. It was drawn as a 200-session sample, of which 17 were withheld during pre-release screening: 15 after a review of embedded images and 2 after an identifier scan. The full subset of 26,999 sessions is released on acceptance, once the redaction audit is complete. The paper's appendix… See the full description on the dataset page: https://huggingface.co/datasets/submit2331/ACT-2.5B.tabulartext-generationn<1K0 likes75 downloads15d agoHugging Face13egygi /computer-use-large-actions computer-use-large-actions 9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video). Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0. Split by software category examples vscode 2,500 autocad 2,500 blender 1,000 excel 1,000 photoshop 1,000 salesforce 1,000 VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.tabulartext-generation1K<n<10K0 likes58 downloads25d agoHugging Face14values-md /when-agents-act Dataset Card for "When Agents Act" Dataset Summary This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute). Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.tabulartext-classificationn<1K1 likes52 downloads10mo agoHugging Face15iknow-lab /JudgeBias-DPO-RefFree-LOSO-action JudgeBias-DPO-RefFree-LOSO-action A Leave-One-Swap-Type-Out (LOSO) variant of JudgeBias-DPO-RefFree for evaluating out-of-distribution generalization of DPO-trained LLM judges. Held-out Swap Type Field Value Swap type Action Axis τ⁻ (error) Removed dataset action_antonym_100pct Description action antonym substitution All DPO pairs derived from action_antonym_100pct have been removed from both training and validation splits. The model trained… See the full description on the dataset page: https://huggingface.co/datasets/iknow-lab/JudgeBias-DPO-RefFree-LOSO-action.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face16actixon /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/actixon/US-PD-Books.tabulartext-generation100K<n<1M0 likes16 downloads6mo agoHugging Face17fineset-io /vision-language-action-papers Vision-Language-Action (VLA) & Robot Learning Papers — FineSet A research-paper dataset on Vision-Language-Action (VLA) & Robot Learning Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Vision-Language-Action (VLA) & Robot Learning Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/vision-language-action-papers.tabulartext-classificationn<1K0 likes16 downloads3mo agoHugging Face18louisbrulenaudet /code-action-sociale-familles Code de l'action sociale et des familles, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-action-sociale-familles.tabulartext-generation1K<n<10K0 likes15 downloads1y agoHugging Face19PhillyMac /Active_Listening_Content_1 Active-Listening-Content-1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Active_Listening_Content_1.tabulartext-generation1K<n<10K0 likes14 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.