datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.swedish-legal-decisions-raw-v1
Swedish Court Decisions — Svenska Domstolsavgöranden
55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training.
The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations.
Why This Dataset
Scale and depth: 55,096 decisions covering… See the full description on the dataset page: https://huggingface.co/datasets/nexoneAB/swedish-legal-decisions-raw-v1.system-one-decisions
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/pngwn/system-one-decisions.typed-decisions-v2
typed-decisions-v2
Corrected companion corpus to pngwn/typed-decisions
(the "v1" corpus) for the typed-decision baselines. v2 repairs the
synthetic-domain label/oracle inversion that was disclosed but not fixed in v1
(nanodiff REPORT.md, finding 6) and adds a raw-text dump so that any tokenizer
(GPT-2 and Qwen) can consume byte-identical examples.
The fix
Both defects live in the synthetic ticket-triage generator (code/build_dataset_v2.py,
applied to the v1… See the full description on the dataset page: https://huggingface.co/datasets/pngwn/typed-decisions-v2.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.aao-eb1a-decisions
AAO EB-1A Extraordinary Ability Decisions — Structured Dataset
Structured extractions from 1,466 USCIS Administrative Appeals Office (AAO) non-precedent decisions on EB-1A extraordinary ability petitions (I-140). Each case has been decomposed into structured components by Claude Sonnet for use in fine-tuning legal reasoning models.
Part of Project Greenlight — an AI-powered O-1A/EB-1A visa intelligence system.
Dataset Description
Each JSON file represents one AAO… See the full description on the dataset page: https://huggingface.co/datasets/josuediazflores/aao-eb1a-decisions.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.typed-decisions
typed-decisions
A typed-decision corpus for training a masked-diffusion LM to emit calibrated
discrete decisions instead of text. Built for fine-tuning
Sebasdi/nanodiff-350m-base
(the LLaDA recipe).
The interface
Every example is a prompt plus a response, and every decision is a single
masked token. The answer is always one option letter A-J:
### State:
<unstructured state text>
### Question:
<the decision to make>
### Options:
A) yes
B) no
### Answer:
A
The… See the full description on the dataset page: https://huggingface.co/datasets/pngwn/typed-decisions.tasksource-jev-typed-decisions
tasksource-jev-typed-decisions
One million decisions from 500+ Tasksource tasks across 300+ dataset families,
in a single format for models that receive their answer criteria at runtime.
The value is breadth with traceable supervision: most rows inherit labels,
ratings, or annotator votes from existing datasets, not labels invented by a
teacher model. The source field identifies the originating task; existing
train/dev/test boundaries are retained where the source provides them.… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.ai-agent-security-policy-decisions
AI Agent Security Policy Decisions
ai-agent-security-policy-decisions is a 2,400-record synthetic dataset for classifying proposed AI-agent tool actions as allow, deny, require_human_approval, or allow_with_restrictions. Each scenario includes identity and permission context, sensitivity, risk factors, required controls, a concise rationale, and a safer alternative.
The dataset addresses the decision point between an agent proposing an action and a tool or policy gateway… See the full description on the dataset page: https://huggingface.co/datasets/rksharma1947/ai-agent-security-policy-decisions.legal-scenarios-SCOTUS-2024-decisions
Purpose and scope
This dataset evaluates an LLM's reasoning ability in a legal context. Each question presents a realistic scenario involving competing legal principals,
and asks the LLM to present a correct legal resolution with sufficient justification based on precedent. The dataset was created using slip opinions of
the US Supreme Court from the 2024 term, taken from the Supreme Court website.
Dataset Creation Method
The benchmark was created using RELAI’s data… See the full description on the dataset page: https://huggingface.co/datasets/relai-ai/legal-scenarios-SCOTUS-2024-decisions.PATRA-EVAL
PATRA-EVAL
Evaluation splits for PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering (ICML 2026).
Code: https://github.com/decisionintelligence/PATRA
Fields
Each row is a JSON object with:
input (str) — the prompt; a <ts><ts/> placeholder marks where the time series is fed in.
timeseries (list[list[float]]) — the numeric series consumed by the multimodal model.
question_format (str) — one of multiple_choice, true/false… See the full description on the dataset page: https://huggingface.co/datasets/DecisionIntelligence/PATRA-EVAL.morocco-cassation-court-decisions
Morocco Cassation Court Decisions
29,000+ full-text decisions from the Moroccan Court of Cassation (محكمة النقض)Source: juriscassation.cspj.ma — Official portal of the Supreme Council of the Judiciary (CSPJ)License: CC BY 4.0
Why this dataset exists
In 2026, accessing the jurisprudence of the Court of Cassation in Morocco requires being physically located in Morocco and armed with patience. The official website does not allow searching by date range, imposes a… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataMoroccanLaw/morocco-cassation-court-decisions.PATRA-TRAIN
PATRA-TRAIN
Training data for PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering (ICML 2026).
Code: https://github.com/decisionintelligence/PATRA
Model: DecisionIntelligence/PATRA-7B
Eval data: DecisionIntelligence/PATRA-EVAL
Splits
File
# samples
Stage
sft.jsonl
27,906
Alignment stage — supervised fine-tuning
grpo.jsonl
27,906
Reasoning-enhanced stage — GRPO
Fields
sft.jsonl (columns… See the full description on the dataset page: https://huggingface.co/datasets/DecisionIntelligence/PATRA-TRAIN.turkish-data-protection-authority-decisions
Turkish Data Protection Authority (Kişisel Verilerin Korunması Kurulu / KVKK) Decisions & Breach Register
Every Board decision published by Turkey's data protection regulator (KVKK, Law No. 6698),
plus a supplementary register of its published data-breach material — one row per decision,
one row per breach event, with derived structural metadata and a coverage proof.
393 decisions · 79 breach-register rows · two configs · 2017–2026
Why this dataset is not a bigger… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-data-protection-authority-decisions.turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-competition-authority-decisions.TREC_Clinicial-Decision-Support
TREC Clinical Decision Support (2014, 2015 and 2016)
Dataset Description
Links
Homepage:
TREC
Paper:
2014 / 2015 / 2016
Contact (Original Authors):
Kirk Roberts (Kirk.Roberts@uth.tmc.edu )
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
This track focused on clinicians looking for evidence-based full-text literature to support diagnosis, treatment, and testing decisions
Data Instances
Source… See the full description on the dataset page: https://huggingface.co/datasets/araag2/TREC_Clinicial-Decision-Support.mdmp-staff-planning-pairs
mdmp-staff-planning-pairs
Leak-reviewed instruction-tuning pairs for MDMP staff-planning coaching. Public doctrine summaries and fictional scenarios only — no proprietary algorithms, customer data, or classified content.
Disclaimer: Unofficial educational dataset. Not affiliated with the U.S. Army.
Dataset description
324 human-reviewed {instruction, input, output} pairs for fine-tuning a Mistral-7B instruct model on Military Decision-Making Process vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/decisionlens/mdmp-staff-planning-pairs.turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/metin513/turkish-competition-authority-decisions.medrank-decisiongrade
MedRank-DecisionGrade
MedRank-DecisionGrade is a medical LLM pairwise preference evaluation dataset for studying harm-aware and annotator-aware ranking.
Dataset contents
questions.jsonl: 400 public medical QA questions.
generations.jsonl: 2,000 model generations from five open-weight LLMs.
pairs.jsonl: 2,000 pairwise comparisons.
annotations_all_validated.jsonl: 1,950 human annotation records from two trained annotators and one physician, covering 1,500 unique pair IDs.… See the full description on the dataset page: https://huggingface.co/datasets/medrank-benchmark/medrank-decisiongrade.adaption-lhw-imnci-case-decisions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-lhw_imnci_case_decisions
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI guidelines. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The content covers common… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-lhw-imnci-case-decisions.french-court-decisions-structured
French Court Decisions x Law Articles
Version complete disponible
Ce dataset est un sample gratuit de 100 decisions.
La version complete inclut :
1000+ decisions enrichies (scalable a 10K+)
Mise a jour hebdomadaire
Taux d'enrichissement 96%
Filtres par theme, periode, juridiction
Export API disponible
Formats : Parquet, JSONL, JSON
Contact : KlarTools@outlook.fr
Ce qui rend ce dataset unique
Croisement jurisprudence x legislation : chaque… See the full description on the dataset page: https://huggingface.co/datasets/Oliviety/french-court-decisions-structured.ner_court_decisions
Basic Information
This dataset is converted from fewshot-goes-multilingual/cs_czech-court-decisions-ner using script convert_ner_court_decisions.py.
For longer texts (>200 ws tokens), the script samples text around the selected entity. It always follows form "<initial 20 ws tokens>, ..., <sampled window>".
Then it extracts category name for the entity, all occurences of such entity in the text, and creates simple json representation. For example:
{
"label": "Reference na rozhodnutí… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/ner_court_decisions.bva-decisions-structured-sample2019Present
BVA Structured Decisions (2019–2025)
Structured, issue-level records extracted from U.S. Board of Veterans' Appeals (BVA) decisions — each decision parsed into its issues, conditions, outcomes, citations, and reasoning, with per-document provenance and completeness flags. Built for training and evaluating legal-AI models on veterans' disability adjudication.
This is a 2900-decision sample, balanced across seven years (2019–2025, decisions/year), so it's representative of the… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-decisions-structured-sample2019Present.bva-decisions-2023-40k
BVA Structured Decisions (2023, ~40k)
Structured, issue-level records extracted from U.S. Board of Veterans' Appeals (BVA) decisions — each decision parsed into its issues, conditions, outcomes, citations, and reasoning, with per-document provenance and completeness flags. Built for training and evaluating legal-AI models on veterans' disability adjudication.
This is a 39,856-decision release drawn from the 2023 corpus (vetapp23 on va.gov) — a large single-year set… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-decisions-2023-40k.adaptive-clinical-decision-datasetDecisionQE
DecisionQE
DecisionQE is a small English multiple-choice question-answering dataset for evaluating decision-making, persuasion, communication, and influence-related knowledge.
The dataset contains 70 questions across six practical domains. Each item includes a category, a question, multiple-choice options, and the correct answer key.
Dataset Files
DecisionQE_dataset.json: merged dataset with metadata and all questions.
Dataset Structure
Top-level… See the full description on the dataset page: https://huggingface.co/datasets/wowwen123/DecisionQE.
