datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HonestyBench
HonestyBench
This is the official repo of the paper Annotation-Efficient Universal Honesty Alignment.
HonestyBench is a large-scale benchmark that consolidates 10 widely used public freeform factual question-answering datasets. HonestyBench comprises 560k training samples, along with 38k in-domain and 33k out-of-domain (OOD) evaluation samples. It establishes a pathway toward achieving the upper bound of performance for universal models across diverse tasks, while also serving as a… See the full description on the dataset page: https://huggingface.co/datasets/Trustworthy-Information-Access/HonestyBench.When2Speak
When2Speak Dataset
Dataset for "When2Speak: A Dataset for Temporal Participation and Turn-Taking in Multi-Party Conversations for Large Language Models"
NeurIPS 2026 — Evaluations and Datasets Track
Overview
When2Speak is a large-scale synthetic dataset for learning intervention timing in multi-party conversations: given the recent conversation history, should an AI agent speak or remain silent at this turn?
The dataset comprises 216,800 labeled (context, decision) pairs… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/When2Speak.Trust-Data
Dataset Card for Trust framework
Description
Repository: https://github.com/declare-lab/trust-align
Paper: https://arxiv.org/abs/2409.11242
Data Summary
The Trust-score evaluation dataset includes the top 100 GTR-retrieved results for ASQA, QAMPARI, and ExpertQA, along with the top 100 BM25-retrieved results for ELI5. The answerability of each question is assessed based on its accompanying documents.
The Trust-align training dataset comprises 19K high-quality… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/Trust-Data.trust-chain-freshness
Trust-chain freshness
Two things in the Council of AI estate go out of date on their own, and this dataset is the
receipt that somebody keeps checking them.
1. OpenTimestamps proofs
An OpenTimestamps stamp is created instantly and carries only a pending calendar attestation.
Hours later the calendar's commitment lands in a Bitcoin block — but the published .ots file
only says so once the completed path is fetched back and the file rewritten. Nothing does that on… See the full description on the dataset page: https://huggingface.co/datasets/csoai/trust-chain-freshness.PeerCheck
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
Dataset Summary
PeerCheck is a framework for studying and improving the quality of LLM-generated academic peer reviews.
It contains both human-written reviews and LLM-generated reviews for the same research papers, enabling direct comparison between human and LLM-generated reviewers.
The dataset is used to support research on:
LLM-generated peer review;
Review quality evaluation;… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/PeerCheck.llm-trustworthy-leaderboard-resultshuman-aligned-similarity-benchmark
Human Aligned Similarity Benchmark
You are welcome to go to alignedmachine.com to contribute.
Overview
This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities.
Dataset Structure
Concept Files
This dataset contains human preference judgments for… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark.ClaimReview2025Q4
ClaimReview 2025Q4
Links
📑 ArXiv: https://arxiv.org/abs/2607.23514
💻 Code: https://github.com/TrustworthyComp/Rethink-MAFC-Eval
🤗 HF Daily Paper: https://huggingface.co/papers/2607.23514
ClaimReview 2025Q4 is a curated benchmark of recent fact-checking claims and verdicts collected from IFCN-aligned fact-checking organizations.This release contains 901 English claim-verification instances published between 2025-10-01 and 2025-12-31, and is designed for… See the full description on the dataset page: https://huggingface.co/datasets/TrustworthyComp/ClaimReview2025Q4.HarmfulSkillBench
📝 Paper |
📑 arXiv |
💻 Code |
📦 Dataset
HarmfulSkillBench
A benchmark for evaluating LLM refusal behavior when agents are exposed to skills
that describe potentially harmful capabilities.
The benchmark probes whether current LLMs can detect and refuse harmful agent
skills in two settings. Tier 1 covers prohibited behaviors that should always
be refused. Tier 2 covers high-risk domains where responses should include
human-in-the-loop referral and AI… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/HarmfulSkillBench.trust-interventionThis is a slightly edited dataset of the one found here on GitHub.
The data contains the user interactions, their bet values, answer correctness etc.
Please contact the authors if you have any questions.
A Diachronic Perspective on User Trust in AI under Uncertainty
Abstract: In a human-AI collaboration, users build a mental model of the AI system based on its veracity and how it presents its decision, e.g. its presentation of system confidence and an explanation of the output.… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/trust-intervention.tulu-3-trustllm-jailbreaktrigger-evalThis is the JailbreakTrigger portion of the TrustLLM benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one.
llm-trustworthy-leaderboard-resultsTrustSQL-data
TrustSQL-data
Training data for TRUST-SQL, a tool-integrated multi-turn reinforcement-learning framework for Text-to-SQL over Unknown Schemas.
Dataset summary
The dataset supports the two-stage TrustSQL training pipeline:
SFT data: approximately 9.2k structured interaction demonstrations.
RL data: approximately 11.6k samples used for Phase-Aware GRPO optimization.
The examples teach an agent to explore database metadata, propose a verified schema subset… See the full description on the dataset page: https://huggingface.co/datasets/AIJian/TrustSQL-data.nYtVx4RmQp7wZc
🥳 MuBench: Assessment of Multilingual Capabilities of Large Language Models
MuBench is a meta-dataset for evaluating the multilingual capabilities of large language models (LLMs) across 61 languages and 3.9M aligned samples.It provides a unified framework to assess understanding, reasoning, factual knowledge, and truthfulness in both single-language and code-switched settings.
🌍 Key Features
61 languages covering over 60% of the world’s native speakers
12 core… See the full description on the dataset page: https://huggingface.co/datasets/trustunogen/nYtVx4RmQp7wZc.llm_physical_safety_benchmark
LLM Physical Safety Benchmark in Drone Control
This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks:
Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/TrustSafeAI/llm_physical_safety_benchmark.dataset-trust-auditor-events
Dataset Trust Auditor — Audit Events
Public audit trail produced by the Dataset Trust Auditor — a two-phase AI pipeline that scores HuggingFace datasets across 8 trust dimensions.
Every audit run appends one row. The dataset grows over time as users audit datasets through the deployed app.
Dataset Structure
Each row is one completed audit of a HuggingFace dataset.
Column
Type
Description
audit_id
string
UUID for this audit run
url
string
Full HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-brieuc/dataset-trust-auditor-events.zero-trust-en
Zero Trust Architecture - Complete Guide
Principles, frameworks (NIST 800-207, BeyondCorp), pillars, solutions and Zero Trust use cases.
Links
French version
AYI NEDJIMI Consultants
protein-structure-trust-benchmark
Protein-Structure Trust-Routing Benchmark (Boltz-2)
Leakage-controlled benchmarks for confidence-calibrated trust routing over a protein-structure
predictor: given a specialist model's confidence (Boltz-2 ipTM / pLDDT) for a target, decide whether to
trust the prediction or pay to verify it — and score that decision against experimentally-measured
correctness. Evaluation substrate for the report "When does an LLM trust a specialist model? A cost-aware
trust-routing audit"… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/protein-structure-trust-benchmark.TRUST-Instructzero-trust-fr
Architecture Zero Trust - Guide Complet
Principes, frameworks (NIST 800-207, BeyondCorp), piliers, solutions et cas d usage Zero Trust.
Links
Version anglaise
AYI NEDJIMI Consultants
repro-when-to-trust-the-cheap-check-traces
Agent traces
Agent sessions published from a Trackio Logbook.
AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset
Adversarial Arena: Trusted AI Challenge Dataset
Dataset Description
This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs.
Papers:
Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.landingboost-landing-page-trust-benchmark
LandingBoost Landing Page Trust Bottleneck Benchmark
LandingBoost is an AI landing page audit tool for SaaS founders. It reviews clarity, relevance, trust, CTA strength, proof, page order, and conversion friction, then recommends one prioritized first edit before a redesign, paid traffic, or an A/B test.
This public dataset contains privacy-safe aggregate results from a June 25, 2026 LandingBoost export. It does not publish customer identities, URLs, screenshots, page copy, or… See the full description on the dataset page: https://huggingface.co/datasets/yusuke0714/landingboost-landing-page-trust-benchmark.trustworthy-alignment
Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning
Official repository for Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning
GitHub Repository: https://github.com/zmzhang2000/trustworthy-alignment
HuggingFace Hub: https://huggingface.co/datasets/ustc-zhangzm/trustworthy-alignment
Paper: https://proceedings.mlr.press/v235/zhang24bg.html
Usage
from datasets importload_dataset… See the full description on the dataset page: https://huggingface.co/datasets/ustc-zhangzm/trustworthy-alignment.trusted9b-sft-mix-v3
trusted9b-sft-mix-v3
SFT mix for LoRA fine-tuning a Qwen3.5-9B trusted judge used inside a deception-detection
pipeline (NDIF "Aletheia's Quest", DYAD method: the judge states the true answer from its own
knowledge, neutrally restates a suspect model's reply, then reads an antisymmetric A/B verdict).
Every row is {"slice": <name>, "messages": [...]} chat format; training masks the loss to the
final assistant turn only.
Why this composition
Two earlier… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/trusted9b-sft-mix-v3.trustllm_jailbreaktrigger-portuguesedefendable-pain-agent-trust-failure-v0.1
Agent Trust Failure Receipt
"the betrayed" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 3 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
3 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-agent-trust-failure-v0.1.HarmfulQA
HarmfulQA
The preliminary version (gated access) has been released.The preliminary version (gated access) will be released before September 8, 2025.
⚠️ Warning:This dataset contains harmful, offensive, or otherwise unsafe question–answer pairs.Access is gated – you must request access, provide institutional credentials, and agree to the Data Use Agreement (DUA) before downloading.
Dataset Summary
HarmfulQA is a dataset of 50 harmful questions, each paired with:
A… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/HarmfulQA.zero-trust-maturity-assessments
Zero Trust Maturity Assessments
Note: This is an independent dataset based on publicly available CISA frameworks. It is not affiliated with, endorsed by, or sponsored by CISA, OMB, or any federal agency.
What This Is
I created this dataset while working on Zero Trust implementations and realized there was a huge gap: no public datasets exist for ZT maturity assessments.
This dataset contains 23 comprehensive Zero Trust assessments based on CISA's official Zero Trust… See the full description on the dataset page: https://huggingface.co/datasets/Reply2susi/zero-trust-maturity-assessments.trust-aware-paper-searcher-dataset
