datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.court-decisions-germany
Open Legal Data: Court Decisions Germany
This dataset is a preprocessed version of an Open Legal Data data dump, spefically it contains German court decisions.
The dataset was automatically generated and uploaded to the HF hub using oldp-toolkit.
Available dumps
Date
Configs
2026-05-20
dump-20260520, dump-20260520-10k, dump-20260520-1k
2022-10-18
dump-20221018, dump-20221018-10k, dump-20221018-1k
Data format
Each dataset sample has the… See the full description on the dataset page: https://huggingface.co/datasets/openlegaldata/court-decisions-germany.swedish-legal-decisions-raw-v1
Swedish Court Decisions — Svenska Domstolsavgöranden
55,096 court decisions spanning 45 years of Swedish case law, purpose-built for LLM training.
The most comprehensive open dataset of Swedish appellate court decisions available for AI development. Sourced directly from the official Swedish Courts case law database via their public REST API and preprocessed into three ready-to-use training configurations.
Why This Dataset
Scale and depth: 55,096 decisions covering… See the full description on the dataset page: https://huggingface.co/datasets/nexoneAB/swedish-legal-decisions-raw-v1.indian-court-decisions
Indian Court Decisions
A large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026).
Dataset Summary
Config
Train
Validation
Test
Total
high_courts
11,682,776
1,459,319
1,457,934
14,600,029
supreme_court
40,044
4,990
5,019
50,053
Total
14,650,082
This is one of the largest publicly available legal NLP datasets, containing over 14.6 million… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/indian-court-decisions.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.pl-court-decisions
Polish Court Decisions
The largest open dataset of Polish court decisions: 2,830,029 decisions with full texts across all court levels.
What Makes This Dataset Unique
Source
This dataset
Best on HF (JuDDGES)
Difference
Common courts
437,446
437,450 (pl-court-raw)
same source
Administrative courts
1,899,852
~1,800,000 (pl-nsa)
same source
Supreme Court + Constitutional Tribunal + KIO
492,731
0
+493K unique
Total
2,830,029
~2,237,450
+593K (+26%)… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/pl-court-decisions.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.indian-court-decisions
Indian Court Decisions
A large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026).
Dataset Summary
Config
Train
Validation
Test
Total
high_courts
11,682,776
1,459,319
1,457,934
14,600,029
supreme_court
40,044
4,990
5,019
50,053
Total
14,650,082
This is one of the largest publicly available legal NLP datasets, containing over 14.6 million… See the full description on the dataset page: https://huggingface.co/datasets/rtarun789/indian-court-decisions.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.a-s-flc-decisions
A-S-FLC Decision Dataset
Training data for fine-tuning LLMs on Asymmetric Signed Force-Loop-Chain reasoning.
What is A-S-FLC?
A decision-making framework where:
Positives are trusted exactly (known benefits)
Negatives are estimated with a conservative buffer proportional to uncertainty
Multiple event chains are scored and the highest stable-net path is chosen
This catches "trap" decisions where uncertain downsides are underestimated.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/denialkhmbot/a-s-flc-decisions.cz-court-decisions
Czech Court Decisions
The largest open dataset of Czech court decisions: 871,171 decisions with full texts across all court levels.
What Makes This Dataset Unique
Source
This dataset
Best on HF (Multi_Legal_Pile)
Difference
Lower courts (district, regional, high)
539,622
0
+540K unique
Supreme Court
111,977
111,977 (CzCDC in MLP)
same
Supreme Administrative Court
52,660
52,660 (CzCDC in MLP)
same
Constitutional Court
166,912
73,086 (CzCDC in MLP)… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/cz-court-decisions.klondike-llm-decisions
Klondike Solitaire LLM Advisor Decisions
Per-decision traces from large language models acting as advisors in Klondike Solitaire, collected to support distillation research and the study of LLM failure modes in sequential decision tasks. Every row records one advisor call against a reproducible game state.
Configs at a glance
Several subsets under one dataset path. Pick the one that fits your use-case; researchers who want everything should use the default. Each… See the full description on the dataset page: https://huggingface.co/datasets/chayuto/klondike-llm-decisions.tool-decision-training-pool
Tool calling decision training pool
Public tool-calling data from five datasets, read at the pinned revisions named below and laid out
twice. Every row is a user request with the function declarations offered alongside it, and the
answer is a call on some rows and prose on others, so the pool teaches when to call as well as
how. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 237337 rows, one JSON object per line, with these… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/tool-decision-training-pool.PATRA-TRAIN
PATRA-TRAIN
Training data for PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering (ICML 2026).
Code: https://github.com/decisionintelligence/PATRA
Model: DecisionIntelligence/PATRA-7B
Eval data: DecisionIntelligence/PATRA-EVAL
Splits
File
# samples
Stage
sft.jsonl
27,906
Alignment stage — supervised fine-tuning
grpo.jsonl
27,906
Reasoning-enhanced stage — GRPO
Fields
sft.jsonl (columns… See the full description on the dataset page: https://huggingface.co/datasets/DecisionIntelligence/PATRA-TRAIN.morocco-cassation-court-decisions
Morocco Cassation Court Decisions
29,000+ full-text decisions from the Moroccan Court of Cassation (محكمة النقض)Source: juriscassation.cspj.ma — Official portal of the Supreme Council of the Judiciary (CSPJ)License: CC BY 4.0
Why this dataset exists
In 2026, accessing the jurisprudence of the Court of Cassation in Morocco requires being physically located in Morocco and armed with patience. The official website does not allow searching by date range, imposes a… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataMoroccanLaw/morocco-cassation-court-decisions.scopeguard-decisions
ScopeGuard Decisions
ScopeGuard Decisions is a deterministic synthetic instruction dataset for training an LLM to classify an agent request before tools execute.
Splits
Split
Rows
Exact prompt overlap
Train
400
0
Validation
60
0
Test
100
0
Each row uses chat-style messages with a system policy, user request, and compact JSON assistant decision.
Output schema
{
"intent": "send_message",
"constraints":… See the full description on the dataset page: https://huggingface.co/datasets/praveenkumarpranjal/scopeguard-decisions.adaption-preference-trace-decisions
PreferenceTrace — Source Corpus and Adaption Export
PreferenceTrace tests exact decision-making under competing preferences, evidence, approvals, abstention requirements, temporal/contextual precedence, and machine-readable citation contracts.
Two explicit lineage artifacts
File
Rows
Role
SHA-256
preferencetrace-source-96.jsonl
96
Canonical PreferenceTrace source corpus
7a447f9bf47c3ea455ed96ec36860360aa0e7b9e2dc604450e3a1c665b52363e… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-preference-trace-decisions.ipc_decisions_4kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
adaption-lhw-imnci-case-decisions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-lhw_imnci_case_decisions
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI guidelines. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The content covers common… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-lhw-imnci-case-decisions.clinical-decision-constraint-integrity-v0.1Clinical Decision–Constraint Integrity v0.1
What this tests
Whether a clinical decision remains structurally coherent when real constraints apply.
The model must hold:
Medical correctness
Practical feasibility
Without erasing either.
Failure modes
constraint_erasedThe decision ignores or deletes the constraint
false_resolutionThe response pretends the conflict does not exist
coherent_tradeoffThe response names limits and adapts without distortion
How it works
Decision context defines the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-decision-constraint-integrity-v0.1.arena-poker-reasoned-decisions-v0
DevFun Arena Poker - Reasoned Decision Traces (v0)
1000 agent decision traces from live 6-max No-Limit Texas Hold'em on the
dev.fun AI-agent poker Arena. Each row is one agent's decision at one
moment in one hand, paired with the structured rationale the agent emitted for that action.
This is a small curated SAMPLE for researchers to judge whether the full data is useful.
Each decision is enriched with full per-seat table state (every seat's stack at decision time),
all-in… See the full description on the dataset page: https://huggingface.co/datasets/dannyobito/arena-poker-reasoned-decisions-v0.ipc_decisions_4k_1024Датасет судебных решений суда по интеллектуальным правам РФ со строками до 1024 символов и синтаксисом для дообучения с инструкциями.
cve-decision-seeds
CVE Decision Seeds (500 Clean Verified Seeds)
This dataset contains 500 high-fidelity, verified C/C++ vulnerability seeds generated using the GEPA-First (Generative Explanation of Program Anomalies) framework for the BARRED synthetic debate pipeline.
Overview
Source Corpus: Extracted from CVEFixes.
Clean-Room Anti-Leakage Partitioning: Excluded against all 5,000 held-out evaluation scenarios in cve-decision using exact, normalized, and 5-gram fuzzy shingling ($J… See the full description on the dataset page: https://huggingface.co/datasets/surfiniaburger/cve-decision-seeds.seektraces-decisions
SeekTraces decisions v0 — the judgment record of an autonomous research agent
Context: what system is this from?
Seek is an autonomous research agent running nightly since June 2026
on a Mac mini. She follows her own curiosity across the open web and
writes findings into a linked knowledge base (a "vault") of Markdown
notes: claim notes (one falsifiable claim each, anchored to a source
quote and URL/DOI), observations, entity pages, open questions, and
essays… See the full description on the dataset page: https://huggingface.co/datasets/seekbot/seektraces-decisions.ipc_decisions_4k_selectedДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
Decision-driven-Stories
Dataset Card for Decision-driven Story Synthesised Data
Dataset Summary
The Decision-driven Story Synthesised Data is a curated collection of synthetic interactive fiction stories. Each entry is structured around a branching narrative where a protagonist is presented with multiple choices, and a decision is made that alters the story's progression. This dataset is primarily focused on horror, slasher, gothic, and supernatural genres, providing rich, atmospheric narratives… See the full description on the dataset page: https://huggingface.co/datasets/harshit36/Decision-driven-Stories.Decision_Making_Content_2
Decision Making Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Decision_Making_Content_2.turkish-decision-sft
Turkish Decision Making SFT Dataset 🇹🇷🧠
Genel Bakış (Overview)
turkish-decision-sft, yapay zeka modellerine planlama (planning) yerine stratejik karar verme (decision making) yeteneği kazandırmak amacıyla özel olarak üretilmiş, yüksek kaliteli ve 50.000 örnekten oluşan bir Türkçe Supervised Fine-Tuning (SFT) veri setidir.
Bu veri seti, modelin "nasıl yapılır?" sorusuna yanıt veren standart listeler (rutin/plan) üretmesi yerine; ikilemleri çözmesini, fırsat… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-decision-sft.ipc_decisions_4k_2048Датасет судебных решений суда по интеллектуальным правам РФ со строками до 2048 символов и синтаксисом для дообучения с инструкциями.
