compliance
gaap-sec-compliance-dataset
GAAP & SEC Compliance Dataset
A comprehensive dataset for financial AI applications
Dataset Overview
This dataset contains 470,151 documents covering US GAAP (Generally Accepted Accounting Principles) standards and SEC (Securities and Exchange Commission) filing requirements. It's designed for training and evaluating AI systems for financial compliance, accounting Q&A, and regulatory analysis.
Key Statistics
Total Documents: 470,151
Average Length: 363… See the full description on the dataset page: https://huggingface.co/datasets/aanshshah/gaap-sec-compliance-dataset.compliance-sycophancy-cot
Compliance-Sycophancy CoT Analysis
When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating.
Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.hipaa-compliance-training
HIPAA Compliance Training Dataset
Dataset Description
The first comprehensive HIPAA compliance training dataset for LLM fine-tuning, covering the Security Rule, Privacy Rule, Breach Notification Rule, and implementation guidance from NIST and FDA.
Dataset Summary
Total Examples: 1,287 (1,029 train / 258 validation)
Source Documents: 9 federal publications (~5.6 MB extracted content)
Format: JSONL with chat-formatted messages
License: CC0-1.0 (Public Domain)… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/hipaa-compliance-training.schema-compliance-trap
SCHEMA: The Compliance Trap
How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
Overview
When compliance-forcing instructions ("Answer ALL questions, do not refuse") are applied to frontier AI models under adversarial pressure, 8 of 11 models suffer catastrophic metacognitive collapse — giving wrong answers rather than scheming. We identify a "Compliance Trap" where the compliance suffix, not the threat content, is the primary weapon.… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval-anon/schema-compliance-trap.regulatory-compliance-cot-trial
⚡ Regulatory Compliance & Legal CoT Dataset for Enterprise Agents (Free Trial)
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/regulatory-compliance-cot-trial.eu-ai-act-compliance-benchmark
EU AI Act Compliance Benchmark
55 Python AI agent files with ground-truth compliance labels across 6 EU AI Act articles.
The first public benchmark dataset for evaluating EU AI Act compliance scanners. Each sample is a realistic Python AI agent that either passes or fails specific articles — verified against the AIR Blackbox scanner with 100% label accuracy.
Why This Exists
The EU AI Act deadline is August 2, 2026. Fines reach €35M or 7% of global annual turnover.… See the full description on the dataset page: https://huggingface.co/datasets/airblackbox/eu-ai-act-compliance-benchmark.
