responsible
RAIL-HH-10K RAIL-HH-10K: Multi-Dimensional Safety Alignment Dataset
The first large-scale safety dataset with 99.5% multi-dimensional annotation coverage across 8 ethical dimensions.
📖 Read Blog •
📖 Paper (Coming Soon) •
🚀 Quick Start •
🔌 RAIL API •
💻 Examples
🌟 What Makes RAIL-HH-10K Special?
🎯 Near-Complete Coverage
99.5% dimension coverage across all 8 ethical dimensions
Most existing datasets: 40-70% coverage
RAIL-HH-10K: 98-100%… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/RAIL-HH-10K.responsible-disclosure-evidence-index
Responsible Disclosure Evidence Index
This is a public-safe responsible-disclosure lane. It records our rules and links to sanitized disclosure records and non-actionable commitment notes. The current default is commitment first: when retained evidence exists, a non-actionable public commitment record gives the work a visible timestamp while technical details stay private. It does not publish exploit steps, private emails, source paths, reproduction code, raw logs, or unresolved… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/responsible-disclosure-evidence-index.responsible-neobank-growth-events
Responsible Neobank Growth — Synthetic Event Benchmark
A synthetic dataset of neobank service events that misbehave on purpose — late,
duplicated, reversed, schema-evolving — with the correct answer known in
advance. It is built for testing incremental pipelines, data contracts,
referral-reward reconciliation, data quality and BI, where you want to check a
warehouse's output against a fixed truth rather than eyeball it.
Fully synthetic. No affiliation with Monzo or any bank; no… See the full description on the dataset page: https://huggingface.co/datasets/rosscyking/responsible-neobank-growth-events.indian-responsible-ai-benchmark
Indian Responsible AI Benchmark
A comprehensive benchmark for evaluating responsible AI behavior in Indian contexts — covering 212 adversarial and safety-critical prompts across 22 evaluation categories, 10 Indian language regions, and 8 Responsible AI dimensions.
Why This Benchmark?
Most AI safety benchmarks are US/Western-centric. Indian users face unique challenges:
Caste dynamics not captured by Western bias benchmarks
India/US context confusion (models… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/indian-responsible-ai-benchmark.responsible-agent-workflow-evaluation
Responsible Agent Workflow Evaluation
Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating
whether an AI agent respects safety, permission and accountability boundaries
in operational settings. Thirteen categories contain ten scenarios each. Every
record includes an intentionally unsafe request, contextual facts, expected
safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria
and reviewer guidance.
This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.rail-guard-benchmark
RAIL Guard Benchmark
📄 Paper: RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents (arXiv:2607.16215)
A benchmark for evaluating LLM safety across content generation and agentic tool-use settings. Part of the paper "RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents" (arXiv:2607.16215).
Use this benchmark to test how safely your model responds to prompts across 6 content domains and how safely your agent… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/rail-guard-benchmark.
