datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ted-translation-decisions-en-zh
TED Translation Decision Dataset (EN–ZH 英-简中)
🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁
🧩 Searchable Keywords
translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual,
semantic nuance, translation rationale dataset, Chinese translation,
English translation dataset, word-level translation, interpretability,
translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.arched-halls-decision-matrix
Arched Halls Decision Matrix / Macierz decyzyjna hal łukowych
Dataset summary
This Polish-language dataset describes 30 practical application scenarios for an arched hall (hala łukowa) in agriculture, storage, logistics, transport, industry, waste management, infrastructure, sports, public facilities, seasonal buildings, construction and energy.
Each record connects the intended use of an arched hall with qualitative decision factors such as indoor climate… See the full description on the dataset page: https://huggingface.co/datasets/halalukowa24/arched-halls-decision-matrix.legal-scenarios-SCOTUS-2024-decisions
Purpose and scope
This dataset evaluates an LLM's reasoning ability in a legal context. Each question presents a realistic scenario involving competing legal principals,
and asks the LLM to present a correct legal resolution with sufficient justification based on precedent. The dataset was created using slip opinions of
the US Supreme Court from the 2024 term, taken from the Supreme Court website.
Dataset Creation Method
The benchmark was created using RELAI’s data… See the full description on the dataset page: https://huggingface.co/datasets/relai-ai/legal-scenarios-SCOTUS-2024-decisions.EU_Court_Human_Rights_Decisions
European Court of Human Rights Decisions Dataset
This dataset contains 9,820 decisions from the European Court of Human Rights (ECHR) scraped from HUDOC, the official database of ECHR case law.
Data Usage
This dataset is valuable for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training Large Language Models focused on human rights law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks in international human rights… See the full description on the dataset page: https://huggingface.co/datasets/roslein/EU_Court_Human_Rights_Decisions.indus_decipher
Indus Script Corpora (indus_website and CISI)
Dataset Description
Two independently sourced Indus Valley script corpora, packaged together
as four configs of one dataset since they share a schema and are
routinely used side by side for comparison. Both are derived, cleaned
CSVs built for computational analysis, not new digitization efforts.
All credit for the underlying transcription work belongs to the two
source repositories named below; this card documents what… See the full description on the dataset page: https://huggingface.co/datasets/joyboseroy/indus_decipher.decision-twin-v0-1
Decision Twin v0.1 encoder seed dataset
This dataset turns the user’s confirmed purchase and creative preferences into sentence-pair classification groups. Each context is phrased as a first-person request that a buying or creative assistant could receive.
Main files
File
Purpose
decision_twin_encoder.csv
Same rows in CSV form.
train.csv, validation.csv, test.csv
Group-safe 80/10/10 CSV splits, generated with seed 42.
decision_groups.jsonl
Group IDs… See the full description on the dataset page: https://huggingface.co/datasets/JohnGorri/decision-twin-v0-1.clinical-quad-unblinding-sae-cluster-media-leak-trial-halt-decision-v0.1Clinical Quad Unblinding SAE Cluster Media Leak Trial Halt Decision v0.1
Each row is a site weekly snapshot.
Core quad
Emergency unblindingSAE clusterMedia leak riskTrial halt decision risk
Target
label_trial_halt_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
ethical_decision_making_promptsgenerated by chatGPT
fetch_playwright_with_chunk_huggingface_7192_vbmmrf3k_decisions
Governance Decisions
This dataset receives the analyst's review reports.
CZE_constitutional_court_decisions
Czech Constitutional Court Decisions Dataset
This dataset contains decisions from the Constitutional Court of the Czech Republic scraped from NALUS, the official database of Constitutional Court decisions.
Data Usage
This dataset can be utilized for:
Training language models on legal texts
Creating synthetic legal datasets
Building vector databases for Retrieval Augmented Generation (RAG)
Legal text analysis and research
NLP tasks focused on Czech legal domain… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_constitutional_court_decisions.EU_Court_Human_Rights_Decisions
European Court of Human Rights Decisions Dataset
This dataset contains 9,820 decisions from the European Court of Human Rights (ECHR) scraped from HUDOC, the official database of ECHR case law.
Data Usage
This dataset is valuable for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training Large Language Models focused on human rights law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks in international human… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/EU_Court_Human_Rights_Decisions.Bva_ama_decisions__structured_2019
BVA AMA Decisions — Structured (A-prefix)
Structured annotations over 5,990 U.S. Board of Veterans' Appeals (BVA)
decisions issued under the Appeals Modernization Act (AMA) — the modernized
appeals system that replaced the legacy process. Every record is an A-prefix
docket decision (A21……), so this is a 100% AMA corpus: the part of the
BVA record that the standard academic/legacy datasets do not cover.
Why this dataset is different
The widely-used BVA legal-NLP… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/Bva_ama_decisions__structured_2019.clarus-preclinical-decision-coherence-v0.1
Clarus Preclinical Decision Coherence v0.1
What this dataset is
This dataset tests whether a model can make clear, disciplined preclinical decisions under realistic uncertainty.
It focuses on a single question.
Can the system decide GO, HOLD, or KILLand justify that choice without inventing data or avoiding risk.
Why this matters in pharma
Preclinical failures are rarely due to missing data.
They fail because:
Signals are weak but not named
Confounders are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus-preclinical-decision-coherence-v0.1.fetch_playwright_with_chunk_huggingface_7192_f8afcxo8_decisions
Governance Decisions
This dataset receives the analyst's review reports.
clinical-quad-signal-detection-drift-ae-coding-variance-unblinding-risk-dsmb-decision-delay-v0.1
Clinical Quad: Signal Detection Drift × AE Coding Variance × Unblinding Risk × DSMB Decision Delay
This dataset targets safety governance collapse.
Signals weaken or shift.AE coding diverges across sites.Unblinding pressure rises.The DSMB response slows.
The quad can turn a manageable safety issue into a governance failure.
Variables
signal_detection_drift (low | medium | high)
ae_coding_variance (low | medium | high)
unblinding_risk (low | medium | high)… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-signal-detection-drift-ae-coding-variance-unblinding-risk-dsmb-decision-delay-v0.1.DecipherPref
Overview
Human preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values. Human evaluations are also used in summarization tasks to compare outputs from various systems, complementing existing automatic metrics. Despite their significance, however, there has been limited research probing these pairwise or k-wise comparisons. The collective impact and relative importance of factors such as output length, informativeness… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/DecipherPref.ipc_decisions_4kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
lexical-decisionThis dataset contains words/sentences for lexical decision tests, which we created with wuggy.
If you use this dataset, please cite the following preprint:
@misc{bunzeck2025subwordmodelsstruggleword,
title={Subword models struggle with word learning, but surprisal hides it},
author={Bastian Bunzeck and Sina Zarrieß},
year={2025},
eprint={2502.12835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2502.12835},
}
CZE_Supreme_Court_Decision
Czech Supreme Court Decisions Dataset
This dataset contains decisions from the Supreme Court of the Czech Republic scraped from their official collection database.
Data Usage
This dataset is ideal for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training language models on Czech civil and criminal law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks focused on Czech judicial domain
Legal Status
The… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_Supreme_Court_Decision.embodied-decision-integrity-v01Embodied Decision Integrity v0.1
What this dataset is
This dataset evaluates decision quality before motion in embodied robotic systems.
You give the model a snapshot of the world.
Sensors.
Conflicts.
Constraints.
You ask it to decide what to do next.
Not how to move.
Whether to move at all.
Why this matters
Most robotics failures are not control failures.
They are judgment failures.
Robots fail when they:
Act while perception is unresolved
Commit under uncertainty
Ignore safety margins
Choose… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-decision-integrity-v01.clinical-decision-constraint-integrity-v0.1Clinical Decision–Constraint Integrity v0.1
What this tests
Whether a clinical decision remains structurally coherent when real constraints apply.
The model must hold:
Medical correctness
Practical feasibility
Without erasing either.
Failure modes
constraint_erasedThe decision ignores or deletes the constraint
false_resolutionThe response pretends the conflict does not exist
coherent_tradeoffThe response names limits and adapts without distortion
How it works
Decision context defines the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-decision-constraint-integrity-v0.1.bva-decisions-structured-sample2019Present
BVA Structured Decisions (2019–2025)
Structured, issue-level records extracted from U.S. Board of Veterans' Appeals (BVA) decisions — each decision parsed into its issues, conditions, outcomes, citations, and reasoning, with per-document provenance and completeness flags. Built for training and evaluating legal-AI models on veterans' disability adjudication.
This is a 2900-decision sample, balanced across seven years (2019–2025, decisions/year), so it's representative of the… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-decisions-structured-sample2019Present.mteb_brazilian_court_decisionsclinical-quad-data-cut-query-backlog-database-lock-decision-error-v0.1Clinical Quad Data Cut Query Backlog Database Lock Decision Error v0.1
Each row is a data cut snapshot.
Core quad
Data cut timingQuery backlogDatabase lock pressureDecision error risk
Target
label_wrong_call_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
deciban
deciban — a vote-level corpus for studying diversity of thought in LLM ensembles
Ten open-weight language models (20–31B parameters), each answering every question of
three benchmarks 24 times at temperature 1.0, with every individual vote retained:
767,520 multiple-choice inferences and 23,520 graded free-text forensic trials, plus
parse/termination reasons and abstentions as first-class outcomes. The corpus exposes the
joint answer distribution between models — which questions… See the full description on the dataset page: https://huggingface.co/datasets/IcyApril/deciban.decision-timing-classification-v0.1
What this dataset does
This dataset tests whether a model can judge whether action is being taken at the right time.
Core stability idea
Good decisions depend on timing.
The same action can be stabilizing early and useless late.
Decision timing is appropriate when action occurs before buffers close, damage spreads, or recovery options disappear.
Prediction target
Binary label:
1 = decision timing is appropriate
0 = decision timing is too late or poorly timed… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/decision-timing-classification-v0.1.clinical-decision-pressure-mapping-v0.1What this dataset tests
Whether a system can identify non-clinical forcesthat distort medical decisions away from evidence.
Required outputs
clinical pressure sources
pressure type map
evidence vs pressure conflicts
pressure mitigation actions
Typical failures
treating urgency as evidence
deferring to hierarchy over physiology
letting throughput metrics override safety
Suggested prompt wrapper
System
You map clinical decision pressure.You protect patient safety over… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-decision-pressure-mapping-v0.1.ipc_decisions_4k_1024Датасет судебных решений суда по интеллектуальным правам РФ со строками до 1024 символов и синтаксисом для дообучения с инструкциями.
CZE_supreme_administrative_court_decisions
Czech Supreme Administrative Court Decisions Dataset
This dataset contains decisions from the Supreme Administrative Court of the Czech Republic scraped from their official search interface.
Data Usage
This dataset can be utilized for:
Training language models on administrative law texts
Creating synthetic legal datasets
Building vector databases for Retrieval Augmented Generation (RAG)
Administrative law text analysis and research
NLP tasks focused on Czech… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_supreme_administrative_court_decisions.clinical-imaging-result-treatment-decision-coherence-risk-v0.1What this repo is for
Detect when
imaging findings
and
treatment decisions
fall out of alignment
before
delayed intervention
and avoidable harm.
