datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
typed-decision-bench
Typed Decision Bench v0.3
Built by Blobfish AI. A benchmark for one-pass decision models: 5,387 items, 25 tasks, 5 use-case
suites. Blobfish designed the tasks, wrote the typed questions, framed each one as a decision a business actually
delegates (use case, vertical), drew stratified seeded panels, froze them, and built the scoring, the contamination
tiers and the quality scorecard. The underlying records are drawn from 21 openly licensed public datasets plus one
generator of… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/typed-decision-bench.a-s-flc-decisions
A-S-FLC Decision Dataset
Training data for fine-tuning LLMs on Asymmetric Signed Force-Loop-Chain reasoning.
What is A-S-FLC?
A decision-making framework where:
Positives are trusted exactly (known benefits)
Negatives are estimated with a conservative buffer proportional to uncertainty
Multiple event chains are scored and the highest stable-net path is chosen
This catches "trap" decisions where uncertain downsides are underestimated.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/denialkhmbot/a-s-flc-decisions.this-that-complex-decisions
this-that-complex-decisions
1,710 decisions where the answer follows from a stated policy applied to a state, and where no
single field of that state gives it away.
1,710 questions 19 decision types 40 domains chance rate 0.258
Each row is a state, a question, a closed set of options, and the index of the one option the
policy selects. The answer is determinate: given the state and the policy there is exactly one
correct choice, and it does not depend on anyone's… See the full description on the dataset page: https://huggingface.co/datasets/limberc/this-that-complex-decisions.swiss_leading_decisions
Dataset Card for Swiss Leading Decisions
Dataset Summary
Swiss Leading Decisions is a multilingual, diachronic dataset of 21K Swiss Federal Supreme Court (FSCS) cases. This dataset is part of a challenging text classification task. We also provide additional metadata as the publication year, the law area and the canton of origin per case, to promote robustness and fairness studies on the critical area of legal NLP.
Supported Tasks and Leaderboards
Swiss Leading… See the full description on the dataset page: https://huggingface.co/datasets/rcds/swiss_leading_decisions.huggingface_filesystem_terminal_12679_q7v2m9_triage_decisionscerto-synthetic-decisions
certo — synthetic decision dataset
Synthetic decisions with a known, exact answer distribution, for training and evaluating
calibrated decision models. Each example is generated from a conditional naive-Bayes evidence
world, so the posterior over the answer is computed in closed form — you can grade a model against
the true probabilities (posterior fidelity), not just accuracy.
Part of certo · project page:
https://altslate-labs.github.io/certo/
Schema (JSONL, one… See the full description on the dataset page: https://huggingface.co/datasets/rajpdus/certo-synthetic-decisions.repro-markov-decision-contests-traces
Agent traces
Agent sessions published from a Trackio Logbook.
eikos-decisions
Eikos Decisions
The training data of Eikos-4B and Eikos-27B, open single-pass typed-decision models.
Each row is one typed decision. It has:
a state (the evidence: a ticket, an email thread, a policy, a table, a log…);
a question of type noul (yes/no), choice (one of N options) or score (ordinal levels);
the options, in the canonical order the model sees them;
a probability distribution over the options (target_probs), which is the soft target the models were trained on.
The… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/eikos-decisions.aimultiple-decision-models-browser
AIM Decision Model Browser Benchmark: 10-task sample
This is 10 of the 50 tasks from AIMultiple's decision model browser benchmark. The benchmark compares decision models (also called System One models) with general LLMs on browser tasks: Jev 1.13, Kev-9B, Laya typed-decisions, Gemini 3.8 Flash and GPT-6 Astra. The other 40 tasks are withheld so they can be reused for later runs.
Article: https://aimultiple.com/decision-models
Files
tasks.jsonl has one task per… See the full description on the dataset page: https://huggingface.co/datasets/AIMultiple/aimultiple-decision-models-browser.german-court-decisions
Dataset Card for german-court-decisions
60k judicial decisions in Germany retrieved on January 1, 2024.
Dataset Description
Language(s) (NLP): German
License: MIT
Copyright notice: Automated retrieval of decisions from federal and state databases in Germany is permitted for non-commercial purposes only. As a result, the use of this dataset is permitted for non-commercial purposes only.
Uses
Prediction of verdicts based on statement of facts.
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/SH108/german-court-decisions.tl_6854_gov_decision_u82os289dhhumanoid-context-aware-decision-dataset
Humanoid Context-Aware Decision Dataset
Dataset for training humanoid AI to make decisions
based on environmental and internal context.
Description
Contains contextual parameters and selected actions
to support intelligent reasoning.
File
context_aware_decision_dataset.json
License
MIT
han-humanoid-decision-response-v1
Humanoid Decision Response Dataset
Overview
Dataset ini merekam bagaimana sistem kognitif humanoid
merespons berbagai situasi lingkungan secara real-time.
Cocok untuk:
Autonomous decision modeling
Risk-aware response system
Behavioral AI research
Features
obstacle_proximity_cm
human_presence_confidence
battery_level_percent
mission_priority_level (1-5)
internal_temperature_celsius
system_latency_ms
environmental_risk_index
Target
decision_action… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-decision-response-v1.council-decisions-benchmark
Teranode Council Decisions Benchmark
Anonymous structural telemetry + head-to-head eval comparisons from the Teranode Council — a multi-agent reasoning system for regulated financial advisory.
This dataset is the public-facing complement to the live system at teranode.ai. It captures every model-pin head-to-head comparison Teranode has run during model selection, plus anonymized structural telemetry from every Council deliberation. The dataset is regenerated nightly from the… See the full description on the dataset page: https://huggingface.co/datasets/teranode-ai/council-decisions-benchmark.ua-council-decisions
Ukrainian Municipal Council Decisions — Masthead Identity Extraction
Structured-extraction dataset of 1,075 Ukrainian municipal council decisions (рішення) from
43 local councils (громади / ради) — balanced to exactly 25 decisions per council, each paired with the five identity fields that appear in the
document masthead. The task: given the full text of a single decision, extract its masthead identity.
These are public government records. All personal names in the data are… See the full description on the dataset page: https://huggingface.co/datasets/oshyshatskyi/ua-council-decisions.Mimir_DecisionClassifierhan-humanoid-decision-priority-shifts-v1
Humanoid Decision Priority Shifts (HDPS)
Abstract
HDPS records dynamic priority adjustments
made during multi-objective humanoid tasks.
It enables research on adaptive decision systems
and hierarchical autonomy modeling.
Fields
decision_id
initial_priority_vector
anomaly_trigger
adjusted_priority_vector
decision_latency
outcome_score
Intended Research
Priority reweighting models
Multi-objective optimization
Adaptive decision control… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-decision-priority-shifts-v1.
