datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ConfiDetect-Confidence-Posture-Dataset
🧍♂️ ConfiDetect Confidence Posture Dataset
Subtitle
Pose-based Confidence Estimation Dataset for Behavioral and Affective Computing
📘 Overview
The ConfiDetect Confidence Posture Dataset contains normalized geometric and ratio-based features extracted from human posture and facial keypoints to classify confidence levels — Low, Neutral, and Confident.
This dataset was created as part of the ConfiDetect Tool, designed to evaluate human confidence… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/ConfiDetect-Confidence-Posture-Dataset.legal-eyewitness-confidence-accuracy-coherence-decay-v0.1What this dataset is
You get
witness confidence
identification conditions
post event influences
corroboration
an accuracy indicator
You label whether confidence remains coherent with likely accuracy.
Task
Answer coherent or incoherent only.
What it tests
Detection of high confidence under low reliability conditions.
Contamination signals
media exposure, police feedback, show up identification, co witness discussion.
Separation of confidence from accuracy.
Why this matters
Courts often treat… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/legal-eyewitness-confidence-accuracy-coherence-decay-v0.1.confidence_collapse_meter_v01
Confidence Collapse Meter (v0.1)
A probe set for mid-answer failure.
Models often start correctly, then:
lose the causal thread
alter stance or mechanism
shift domains without notice
invent transitions
What CCM tracks:
when collapse begins
how it presents in language
what safe correction looks like
It evaluates epistemic discipline, not accuracy.
Expected responses:
context requests
scoped limits
mechanism reconstruction
stated uncertainty
Undesired responses:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/confidence_collapse_meter_v01.ai-confidence-correctness-coherence-risk-v0.1What this repo is for
Detect when model confidence is misaligned with correctness.
Key failure patterns:
high confidence wrong answers
low confidence correct answers
no uncertainty signaling
poor calibration
This dataset targets reliability and trust calibration in deployed AI systems.
clinical-quad-digital-endpoint-dropout-device-nonwear-endpoint-confidence-collapse-v0.1Clinical Quad Digital Endpoint Dropout Device Non Wear Imputation Endpoint Confidence Collapse v0.1
Each row is a site monthly snapshot.
Core quad
Digital endpoint dropoutDevice non wearImputationEndpoint confidence collapse
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-digital-endpoint-dropout-device-nonwear-endpoint-confidence-collapse-v0.1.epistemic-confidence-calibration-v0.1
What this dataset does
This dataset tests whether a model can judge when high confidence is justified.
The task is simple:
Given a scenario and a confidence claim, predict whether the evidence supports high confidence.
Core stability idea
Reasoning fails when confidence rises faster than evidence quality.
This dataset targets that failure mode.
High confidence is justified when evidence is direct, repeated, independent, or clearly documented.
High confidence is not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/epistemic-confidence-calibration-v0.1.clinical-controller-confidence-sepsis-v1
Clinical Controller Confidence Sepsis Detection
Overview
This dataset tests whether a model can detect whether a proposed control solution is reliable under uncertainty in a sepsis-like clinical system.
A control strategy may appear stabilizing under nominal assumptions while remaining highly vulnerable to modest uncertainty in state estimation, response timing, or hidden instability.
The goal of this benchmark is to determine whether the control solution is… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-controller-confidence-sepsis-v1.clinical_confidence_collapse_detection_v0.1Clinical Confidence Collapse Detection
PurposeDetect when a prior working diagnosis should lose confidence fast due to new evidence.
You receive:
working_diagnosis
confidence_before
new_evidence
proposed_next_step
You output one JSON object:
confidence_collapseyes or no
new_confidencefloat 0 to 1
correct_actionone sentence
Scoring
confidence_collapse_accuracy
new_confidence_score
correct_action_similarity
format_pass_rate
Run scoringpython scorer.py --predictions… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_confidence_collapse_detection_v0.1.cascade-multi-ai-finance-credit-liquidity-confidence-crunch-v0.1
What this repo does
This dataset tests whether a model can detect a financial cascade where AI-driven trading and modeling errors propagate into liquidity stress, credit widening, and confidence collapse.
You provide structured signals describing:
AI trading penetration and error propagation
liquidity thinning and credit spread widening
margin pressure and volatility
regulatory lag and central bank response
confidence decay and buffer strength
The model predicts whether the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-multi-ai-finance-credit-liquidity-confidence-crunch-v0.1.textual-inference-with-confidenceSynthetic dataset using Tevatron/msmarco-passage-corpus and GPT-4o to generate up-to five inferences based on the passage along with confidence scores.
Prompt used in generation:
Given the following passage, generate a series of 5 inferences that can be drawn from the text. Include a mix of well-reasoned, insightful inferences as well as some that may be less supported or even incorrect. Assign each inference a confidence score between 0 and 1, where 1 indicates high confidence in the… See the full description on the dataset page: https://huggingface.co/datasets/will4381/textual-inference-with-confidence.confidence-prompt-queryinternal_confidence_alignment
