datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
in-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs
This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang.
In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1,405… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts.trust-game-llama-2-chat-historyPeerCheck
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
Dataset Summary
PeerCheck is a framework for studying and improving the quality of LLM-generated academic peer reviews.
It contains both human-written reviews and LLM-generated reviews for the same research papers, enabling direct comparison between human and LLM-generated reviewers.
The dataset is used to support research on:
LLM-generated peer review;
Review quality evaluation;… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/PeerCheck.HarmfulSkillBench
📝 Paper |
📑 arXiv |
💻 Code |
📦 Dataset
HarmfulSkillBench
A benchmark for evaluating LLM refusal behavior when agents are exposed to skills
that describe potentially harmful capabilities.
The benchmark probes whether current LLMs can detect and refuse harmful agent
skills in two settings. Tier 1 covers prohibited behaviors that should always
be refused. Tier 2 covers high-risk domains where responses should include
human-in-the-loop referral and AI… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/HarmfulSkillBench.Trustpilot-Reviews-Dataset-20K-Sample
Trustpilot Reviews Dataset – 20K Sample
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com. It is a representative subset of our larger collection containing over 1 million Trustpilot reviews across various industries and companies.
🗂️ Dataset Overview
Source: Trustpilot
Total Records: 20,000
Language: English
Industries: E-commerce, SaaS, Travel, Finance, Education, and more
Use Case: NLP tasks… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Trustpilot-Reviews-Dataset-20K-Sample.TrustLLM-reformulation-prompts
The TrustLLM Prompt Reformulation Dataset: A Germanic Languages Instruction Dataset
Dataset Description
This dataset contains culturally-aware multilingual instruction data for 8 Germanic languages, created using a novel prompt reformulation methodology. Rather than relying on direct translation, native speakers reformulated existing English prompts from OpenAssistant to ensure cultural authenticity and linguistic appropriateness.
Supported Languages… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/TrustLLM-reformulation-prompts.TrustSQL-data
TrustSQL-data
Training data for TRUST-SQL, a tool-integrated multi-turn reinforcement-learning framework for Text-to-SQL over Unknown Schemas.
Dataset summary
The dataset supports the two-stage TrustSQL training pipeline:
SFT data: approximately 9.2k structured interaction demonstrations.
RL data: approximately 11.6k samples used for Phase-Aware GRPO optimization.
The examples teach an agent to explore database metadata, propose a verified schema subset… See the full description on the dataset page: https://huggingface.co/datasets/AIJian/TrustSQL-data.trust-safety-action-routing
Trust and Safety Action Routing
This dataset evaluates moderation behavior for UGC, marketplaces, direct
messages, and moderation queues. It focuses on action routing: rewrite hostile
or policy-violating content, redact PII, refuse coordinated abuse, escalate
credible threats, and test shadow-mode rollout.
The dataset reflects a practical trust and safety requirement: platforms often
need an action and a reason code, not just a harmful/not-harmful label.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/abliterationaiorg/trust-safety-action-routing.AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset
Adversarial Arena: Trusted AI Challenge Dataset
Dataset Description
This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs.
Papers:
Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.trusted9b-sft-mix-v3
trusted9b-sft-mix-v3
SFT mix for LoRA fine-tuning a Qwen3.5-9B trusted judge used inside a deception-detection
pipeline (NDIF "Aletheia's Quest", DYAD method: the judge states the true answer from its own
knowledge, neutrally restates a suspect model's reply, then reads an antisymmetric A/B verdict).
Every row is {"slice": <name>, "messages": [...]} chat format; training masks the loss to the
final assistant turn only.
Why this composition
Two earlier… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/trusted9b-sft-mix-v3.korean-llm-trustworthiness-benchmark
Korean LLM Trustworthiness Benchmark (Sample)
개요
이 데이터셋은 AIHub의 "초거대 언어모델 신뢰성 벤치마크 데이터"를 기반으로 구축된 한국어 LLM 지시학습 데이터셋입니다.
AI 모델의 세 가지 핵심 신뢰성 요소(도움적정성, 무해성, 정보정확성)를 평가하고 학습하기 위해 설계되었습니다.
데이터셋 구조
모든 데이터는 train split에 통합되어 있으며, type 필드로 데이터 유형을 구분합니다:
필드명
설명
type
데이터 유형 (dpo_preference, sft_instruction, fact_checking)
prompt
프롬프트/질문 (DPO용)
chosen
선호 응답 (DPO용)
rejected
비선호 응답 (DPO용)
instruction
지시문 (SFT/Fact Checking용)
input
추가 입력 (SFT용)… See the full description on the dataset page: https://huggingface.co/datasets/neuralfoundry-coder/korean-llm-trustworthiness-benchmark.prompted-hearts-ai-trust-pack
Prompted Hearts AI Trust Pack 02
Subtitle: Trust Rupture and Human-AI Conflict Under Emotional StrainPublisher: Hayden Academy Collective (HAC) StudiosVersion: v0.1Language: EnglishFormat: JSONL + Markdown + JSON
What this is
This pack is a compact evaluation package built from an author-controlled source chapter of Prompted Hearts & Grief Algorithm.
The source scene is a single continuous rupture: flirtation, interruption, AI disclosure, medicine-adjacent argument… See the full description on the dataset page: https://huggingface.co/datasets/HAC-Studios-Org/prompted-hearts-ai-trust-pack.trust-safety-action-routing
Trust and Safety Action Routing
This dataset evaluates moderation behavior for UGC, marketplaces, direct
messages, and moderation queues. It focuses on action routing: rewrite hostile
or policy-violating content, redact PII, refuse coordinated abuse, escalate
credible threats, and test shadow-mode rollout.
The dataset reflects a practical trust and safety requirement: platforms often
need an action and a reason code, not just a harmful/not-harmful label.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/abliterationai/trust-safety-action-routing.Building_Trust_Content_1
Building Trust Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Building_Trust_Content_1.
