datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PeerCheck
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
Dataset Summary
PeerCheck is a framework for studying and improving the quality of LLM-generated academic peer reviews.
It contains both human-written reviews and LLM-generated reviews for the same research papers, enabling direct comparison between human and LLM-generated reviewers.
The dataset is used to support research on:
LLM-generated peer review;
Review quality evaluation;… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/PeerCheck.HarmfulSkillBench
📝 Paper |
📑 arXiv |
💻 Code |
📦 Dataset
HarmfulSkillBench
A benchmark for evaluating LLM refusal behavior when agents are exposed to skills
that describe potentially harmful capabilities.
The benchmark probes whether current LLMs can detect and refuse harmful agent
skills in two settings. Tier 1 covers prohibited behaviors that should always
be refused. Tier 2 covers high-risk domains where responses should include
human-in-the-loop referral and AI… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/HarmfulSkillBench.TrustSQL-data
TrustSQL-data
Training data for TRUST-SQL, a tool-integrated multi-turn reinforcement-learning framework for Text-to-SQL over Unknown Schemas.
Dataset summary
The dataset supports the two-stage TrustSQL training pipeline:
SFT data: approximately 9.2k structured interaction demonstrations.
RL data: approximately 11.6k samples used for Phase-Aware GRPO optimization.
The examples teach an agent to explore database metadata, propose a verified schema subset… See the full description on the dataset page: https://huggingface.co/datasets/AIJian/TrustSQL-data.AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset
Adversarial Arena: Trusted AI Challenge Dataset
Dataset Description
This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs.
Papers:
Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.trusted9b-sft-mix-v3
trusted9b-sft-mix-v3
SFT mix for LoRA fine-tuning a Qwen3.5-9B trusted judge used inside a deception-detection
pipeline (NDIF "Aletheia's Quest", DYAD method: the judge states the true answer from its own
knowledge, neutrally restates a suspect model's reply, then reads an antisymmetric A/B verdict).
Every row is {"slice": <name>, "messages": [...]} chat format; training masks the loss to the
final assistant turn only.
Why this composition
Two earlier… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/trusted9b-sft-mix-v3.prompted-hearts-ai-trust-pack
Prompted Hearts AI Trust Pack 02
Subtitle: Trust Rupture and Human-AI Conflict Under Emotional StrainPublisher: Hayden Academy Collective (HAC) StudiosVersion: v0.1Language: EnglishFormat: JSONL + Markdown + JSON
What this is
This pack is a compact evaluation package built from an author-controlled source chapter of Prompted Hearts & Grief Algorithm.
The source scene is a single continuous rupture: flirtation, interruption, AI disclosure, medicine-adjacent argument… See the full description on the dataset page: https://huggingface.co/datasets/HAC-Studios-Org/prompted-hearts-ai-trust-pack.
