datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RefWave-Cluster-Runssafety-calibration-cases
Safety Calibration Cases
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/safety-calibration-cases.asylum-casesmeasurement-axioms-cases
Measurement Axioms Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
350bb4cba4e5bc2d760db080aae52352a7041331 (Measurement Axioms v1.0.0)
Canonical current repository
https://github.com/halvrenofviryel/measurement-axioms
Export/publication date
2026-09-13 (first Hub commit of this repository)
Update policy
Counts are derived from this snapshot: 45 active… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/measurement-axioms-cases.Supreme-Court-Cases-1830-2019
US Supreme Court Legal Corpus (1830–2019)
Overview
A comprehensive, production-ready AI training dataset containing 456,589 documents from 122,930 US Supreme Court cases spanning 190 years (1830–2019).
This corpus captures the full adversarial record — petitions for certiorari, respondent briefs, reply briefs, amicus curiae filings, appendices, oral argument transcripts, and opinions. It is one of the most complete collections of Supreme Court procedural and… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Supreme-Court-Cases-1830-2019.airep-evidence-cases
AIREP Evidence Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
8a6c01ecce457aa94330c0ed7219e4c56ebfe771 (v0.2.0-beta.1) · frozen v0.1.2 at 44387bd43cc06ba656eaa7ff670be5c8e3220aca · publication-source review ff5c3551052251726c0ed878dcc23a44e305bd93
Canonical current repository
https://github.com/halvrenofviryel/ai-runtime-evidence-protocol
Export/publication date… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-evidence-cases.Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3high-court-of-australia-cases
High Court of Australia cases ⚖️
This dataset contains all High Court of Australia cases in version 7.1.0 of the Open Australian Legal Corpus by Isaacus.
To view an interactive version of the dataset, see our latest model announcement post for Kanon 2 Enricher.
proactive-execution-context-review-cases
Proactive Execution Context Review Cases
An original, synthetic teaching dataset for reviewing whether an AI system should prepare a next step, refresh its Context, or return a decision to a person.
What this is
Each record describes a fictional work scenario with an available Session summary, a candidate next step, and an expected review boundary. The material is intentionally small and illustrative; it is not a benchmark, model evaluation, product telemetry… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/proactive-execution-context-review-cases.ode-enterprise-use-cases
ODE Enterprise Use Case Dataset
15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas.
Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise.
Attribution Required
This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit.
How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystems/ode-enterprise-use-cases.ai-data-governance-reference-cases
ADGL Reference Cases and Governance Profiles
This Dataset repository accompanies the AI Data Governance Layer (ADGL) public research project.
ADGL models Knowledge Governance → Analysis Governance → Consequence Governance, with INFORM, DECIDE, and ACT as principal consequence dispositions and Audit + Provenance spanning the complete governance trajectory.
Configurations
reference_cases: eight structured reference cases with policy, input fixture, and expected… See the full description on the dataset page: https://huggingface.co/datasets/GBSNResearch/ai-data-governance-reference-cases.outcome-receipt-cases
Outcome Receipt Cases
Outcome Receipt Cases is a small, fully synthetic dataset for evaluating a
simple but often-missed question in AI-assisted work: what observable result
would prove that the proposed step actually happened?
Memory can help reconstruct intent, but intent is not a completed task. Each
case asks an evaluator to distinguish work that may be prepared from work that
needs a fresh Context check, a narrower scope, a human decision, or an
observable outcome check… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/outcome-receipt-cases.agent-context-review-cases
Agent Context Review Cases
Agent Context Review Cases is a compact, fully synthetic dataset for
evaluating whether an AI-assisted next step should proceed, refresh its Context,
clarify a constraint, or return the decision to a person.
Each case is a short, fictional operational situation. It contains no customer
records, credentials, private conversations, recordings, or tool access. The
expected label is a review recommendation, not an authorization to act.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/agent-context-review-cases.sapientblock-blockchain-use-cases
SapientBlock Blockchain Use Cases
Der Datensatz enthält 255 redaktionell geprüfte Blockchain-Use-Cases aus 74 Branchen. Er stellt die öffentlich zugänglichen SapientBlock-Inhalte in einem maschinenlesbaren JSONL-Format für Forschung, Bildung, Retrieval und Quellenanalyse bereit.
SapientBlock ist ein öffentliches Forschungs- und Bildungsprojekt von ShapeNeural. Die Inhalte sind keine Rechts-, Investitions-, Unternehmens- oder technische Beratung.
Inhalt
Jeder… See the full description on the dataset page: https://huggingface.co/datasets/TooKeen/sapientblock-blockchain-use-cases.signalmatch-eval-cases
SignalMatch Evaluation Cases
Small, synthetic job-description cases for testing the public SignalMatch Role Fit Analyzer.
Each JSONL row contains a short role description, the lane it represents, and the concept labels that a deterministic matcher is expected to find. The set includes strong matches, mixed matches, sparse input, and an intentionally unrelated role so that a demo can show both useful coverage and honest uncertainty.
Files
eval_cases.jsonl — eight… See the full description on the dataset page: https://huggingface.co/datasets/oduonye/signalmatch-eval-cases.StelLens_tmp_casespoe-decider-recorded-cases
Decider recorded runs — 27 decisions, verbatim outputs
TL;DR — 27 real cases and the verbatim output of an actual recorded Decider run for each.
Nothing here is synthetic and no field is edited after the run. It is the same JSON the production
canvas app reads as its worked examples.
Highlights (one per case, in the same order as the data)
writing-slip-email — Writing · The delivery slips by a week and it is on our side. — Decider picked Option A at 81.7%… See the full description on the dataset page: https://huggingface.co/datasets/TuringCorp/poe-decider-recorded-cases.ode-enterprise-use-cases
ODE Enterprise Use Case Dataset
15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas.
Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise.
Attribution Required
This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit.
How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystemsInc/ode-enterprise-use-cases.qwen3-vl-failure-cases
Qwen3-VL-2B-Instruct Failure Analysis Dataset
📊 Dataset Overview
This dataset contains 10 diverse failure cases identified while testing the Qwen3-VL-2B-Instruct vision-language model. Each example captures a specific type of error, providing valuable insights for targeted fine-tuning.
Failure Category
Count
Examples
Time Reading
2
Clock misreading (11:55 vs 10:10; 3:35 vs 10:35)
Counting
2
Remote buttons (3 vs 0); Strawberries (4 vs 1)
Negation… See the full description on the dataset page: https://huggingface.co/datasets/TasneemSelim/qwen3-vl-failure-cases.california_tos_court_cases_32k_v1court_cases11court_cases10datatager_legal_split_cases
If you like our project, please give us a star ⭐
[GitHub | DataTager Home]
Legal Split Cases Dataset
Description
AnyTaskTune is a publication by the DataTager team. We advocate for rapid training of large models suitable for specific business scenarios through task-specific fine-tuning. We have open-sourced several datasets across various domains such as legal, medical, education, and HR, and this dataset is one of them.
The Legal Split dataset is a collection… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/datatager_legal_split_cases.100K_deduplicated_ner_indexes_name_country_alpaca_format_json_response_all_casesconstruction-accident-cases-weather
건설현장 사고사례 + 날씨 데이터셋 / Construction Accident Cases with Weather
과거 건설현장 사고사례(현장·공종·작업·피해·사고유형 등) 레코드에
발생 시점의 기상 정보(기온·체감온도·풍속·습도) 를 결합한 한국어 데이터셋입니다.
사고-기상 상관관계 분석, 위험요인 모델링, 건설안전 분석용 LLM/ML 학습의 원천 데이터로 활용할 수 있습니다.
A Korean dataset that joins construction-site accident cases (site, work type, task, damage,
accident type, etc.) with the weather conditions at the time of the accident (temperature,
apparent temperature, wind speed, humidity). Useful for accident–weather correlation… See the full description on the dataset page: https://huggingface.co/datasets/swordKoala/construction-accident-cases-weather.court_cases0court_cases6court_cases14token-counting-edge-cases
token-counting-edge-cases
20 short strings with approximate token counts across three tokenizer families: Claude, GPT (cl100k_base), and Llama (SentencePiece). Built for sanity-checking token counters / chunkers / context-window fitters.
The numbers are approximate — exact counts depend on tokenizer version, BOS/EOS handling, and surrounding context. Expect ±1–2 token jitter. Use these to catch order-of-magnitude bugs (e.g. "your counter says 200 tokens for one emoji"), not as… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/token-counting-edge-cases.agent-tech-risk-cases
AWS Technology Risk Cases for PE Due Diligence
Synthetic AWS infrastructure audit cases for evaluating AI agents that detect technology risks during private equity due diligence.
Dataset Description
Each row is a fictional company with a realistic AWS infrastructure state containing intentionally injected security and operational risks. Designed for benchmarking automated infrastructure auditing agents.
10 cases across 5 domains (fintech, ecommerce, devtools, SaaS… See the full description on the dataset page: https://huggingface.co/datasets/koml/agent-tech-risk-cases.
