datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-as-policy
Agent as Policy — Real Dual-Arm LLM-Agent Manipulation Trials
Paper · Project page · Code
Summary
162 real-robot trials in which an LLM agent acts directly as the policy on a
bimanual YAM arm setup: it reads camera observations through a tool interface,
issues Cartesian and joint commands, and judges its own completion. Ten
manipulation tasks (block stacking, die flipping, towel folding, part insertion
and assembly, throwing), six models, three reasoning-effort… See the full description on the dataset page: https://huggingface.co/datasets/Agent-as-Policy/agent-as-policy.policy-alignment-verification-dataset
Policy Alignment Verification Dataset
🌐 NAVI's Ecosystem 🌐
🌍 NAVI Platform – Dive into NAVI's full capabilities and explore how it ensures policy alignment and compliance.
🤗 NAVI-small-preview – Access the open-weights version of NAVI designed for policy verification.
📜 API Docs – Your starting point for integrating NAVI into your applications.
📝 Blogpost: Policy-Driven Safeguards Comparison – A deep dive into the challenges and solutions NAVI addresses.
✨… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/policy-alignment-verification-dataset.flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_3COMPASS-Policy-Alignment-Testbed-Dataset
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings.
What is COMPASS?
COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.weak-policy-rolloutWeak policy rollouts
tulu-3-wildchat-reused-on-policy-8b
Llama 3.1 Tulu 3 Wildchat reused (on-policy 8B)
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This preference dataset is part of our Tulu 3 preference mixture:
it contains prompts from WildChat and it contains 17,207 generation pairs (some of which on-policy completions from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-wildchat-reused-on-policy-8b.flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_17optim_policy_pretrain-pythia-160m_lr0.0001_bs24_wp1_wd0.01_ep0_cp35k-mergedAXXXX_jssp_policy_step_train_dispatch_v1policyqa-dutch
PolicyQA-Dutch (MTEB retrieval format)
Dutch government policy-FAQ retrieval task. Given a Dutch policy question, retrieve the relevant policy passage from the corpus.
Reformatted into MTEB retrieval format from beery/PolicyQA-Dutch.
adaptive-policy-v2.8-withreason
PolicyShiftBench
Project Page | Paper | Code
PolicyShiftBench is a comprehensive benchmark designed to evaluate policy-adaptive image guardrailing. Instead of treating safety as an intrinsic property of an image, PolicyShiftBench tests whether a model can decide whether an image violates the currently supplied policy and generalize to held-out policy definitions.
The benchmark features 2,000 policy-discriminative instances over 265 images, where each image is paired with… See the full description on the dataset page: https://huggingface.co/datasets/PolicyShiftGuard/adaptive-policy-v2.8-withreason.tulu-3-wildchat-if-on-policy-8b
Llama 3.1 Tulu 3 Wildchat IF (on-policy 8b)
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This preference dataset is part of our Tulu 3 preference mixture:
it contains prompts from WildChat, which include constraints, and it contains 10,792 generation pairs (some of which on-policy from allenai/Llama-3.1-Tulu-3-8B)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-wildchat-if-on-policy-8b.policy_qa
Dataset for the PolicyQA task in the PrivacyGLUE dataset
task683_online_privacy_policy_text_purpose_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task683_online_privacy_policy_text_purpose_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task683_online_privacy_policy_text_purpose_answer_generation.policystrategies-archive
🏛️ Open-Source Macro-Strategy, Financial History & Intelligence Archive
🌐 Overview & Institutional Mission
This public repository serves as the official open-source knowledge graph and metadata registry for r/policystrategies.
We aggregate, document, and cross-reference declassified historical intelligence dossiers, sovereign debt crises, systemic market manipulations, and geoeconomic conflicts using verified open-source intelligence (OSINT) and primary… See the full description on the dataset page: https://huggingface.co/datasets/stratigahq/policystrategies-archive.EleutherAI_pythia-1b-deduped__dpo_on_policy__tldr
Dataset Card for "EleutherAI_pythia-1b-deduped__dpo_on_policy__tldr"
More Information needed
govuk-policy-qa-pairsThis is a dataset of synthetically generated question and answer pairs on UK government policy papers.
It comes in 2 parts:
Plain text UK government policy papers, scraped from the Gov.uk Policy papers and consultations page. These are in results.json
A series of question and answer pairs on chunk of the above documents, generated using llama_index.finetuning.generate_qa_embedding_pairs and OpenAI GPT3.5 Turbo.
flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_6gemma-4-31B-on-policy-600ktask682_online_privacy_policy_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task682_online_privacy_policy_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task682_online_privacy_policy_text_classification.us-congress-bill-policy-115_117
Dataset Card for us-bills-115_117
All bills introduced to the US House and Senate congress 115, 116, and 117 (2017-2023).
Dataset Details
Dataset Description
This dataset includes all bills introduced to the United States Congress during 2017-2023 (approx. 48k bills).
Included fields:
b_id: Unique string identified
congress: int, the congress designation
title: str, the display title of the bill
summary: str, the earliest dated available summary of the bill… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/us-congress-bill-policy-115_117.VIOLA
VIOLA
Dataset Description
VIOLA is a curated benchmark for evaluating whether LLM-based agents comply with
explicit behavioral policies in a multi-agent pipeline. Each example is a single
agent execution trace — a real run of the CUGA multi-agent system on
AppWorld tasks — in which the target
agent's system prompt was modified to induce a specific policy violation using the
contrary instruction injection pattern: rather than removing a policy, the
distorter replaces it… See the full description on the dataset page: https://huggingface.co/datasets/policy-violation-benchmark/VIOLA.fingpt_convfinqa_sup_sample_from_policy_v1.1_dpo_val_chunk_19fingpt_convfinqa_sup_sample_from_policy_v1.1_dpo_train_chunk_25policy-rag-corpus-metadata
Policy RAG Corpus Metadata (No Raw Data)
This repository is a metadata-only companion for the Policy RAG project built for the Quantic MSSE AI Engineering program.
It does not include the actual PDF files. The source PDFs are hosted in the companion GitHub repository.
What this repo includes
metadata.csv: structured metadata for 11 policy documents (filename, title, category, page count, source type, description)
Citation and provenance notes for reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/mihai-chindris/policy-rag-corpus-metadata.ai-agent-security-policy-decisions
AI Agent Security Policy Decisions
ai-agent-security-policy-decisions is a 2,400-record synthetic dataset for classifying proposed AI-agent tool actions as allow, deny, require_human_approval, or allow_with_restrictions. Each scenario includes identity and permission context, sensitivity, risk factors, required controls, a concise rationale, and a safer alternative.
The dataset addresses the decision point between an agent proposing an action and a tool or policy gateway… See the full description on the dataset page: https://huggingface.co/datasets/rksharma1947/ai-agent-security-policy-decisions.kitchen-object-grasping-manipulation-policy-training-next-pack-5534d90d-f7cb0d04
Cluttered Home Kitchens for Cooking-Assistant Robot
Training dataset of richly cluttered, realistic home kitchens to train a YOLOv8-based cooking-assistant robot that navigates the kitchen, detects and segments objects, distinguishes food from non-food, identifies what needs cleaning, avoids obstacles and manipulates objects. Renders are 640x640 with metric depth, world-space normals (OpenGL linear), albedo and material-index passes, per-frame annotations, and midday lighting… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/kitchen-object-grasping-manipulation-policy-training-next-pack-5534d90d-f7cb0d04.grpo-qwen1.5b-textworld-policy-logitsai-policydrift-2026beauty-brand-policy-data
Beauty Brand Policy Data
This is the 2026-09-16 snapshot of BeautyDeals' 63-row beauty returns and free-shipping comparison for US direct-site shopping. It reproduces the seven fields published in the source table without adding brands, fields, estimates, or independent policy research. Each row preserves the table's own reviewed_on value and the official-source links already attached to that row.
Source comparison: https://lxlex.com/beauty-returns-free-shipping-comparison… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong2387/beauty-brand-policy-data.
