datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airframe-sios-hidden-geometry
Airframe SIOS Hidden Geometry Benchmark
Overview
Airframe SIOS Hidden Geometry is a synthetic relational-reasoning benchmark designed to test whether a system can recover a globally coherent labelled graph when local observations are incomplete, overlapping, conflicting, or actively misleading.
The benchmark is built around two central distinctions:
Metric Evidence ≠ Relational Structure
Local Plausibility ≠ Global Coherence
Each example contains four labelled… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/airframe-sios-hidden-geometry.reasoning-drift-onset-detection-v0.1
Important Evaluation Limitation
Version 0.1 uses a highly regular trajectory structure in which the first drift step is frequently located at Step 4 and visible failure commonly appears at Step 5.
This creates a positional shortcut: a model may achieve inflated onset-detection performance by learning the dataset construction pattern rather than analysing the reasoning trajectory.
Version 0.1 should therefore be treated as a task-definition and scorer-validation release, not as a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-drift-onset-detection-v0.1.clinical-evidence-state-transition-fidelity-v0.1
Clinical Evidence State Transition Fidelity v0.1
A synthetic clinical reasoning dataset for evaluating whether an AI system can update a structured clinical state selectively, proportionately, and consistently when new evidence arrives.
Repository:
ClarusC64/clinical-evidence-state-transition-fidelity-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity
Clinical Evidence State Transition Fidelity v0.1 evaluates whether a model can… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-state-transition-fidelity-v0.1.clinical-evidence-dependency-graph-reasoning-v0.1
Clinical Multi-Evidence State Integration v0.1
Overview
Clinical Multi-Evidence State Integration v0.1 is a structured clinical-reasoning benchmark designed to test whether an AI system can integrate multiple sequential evidence events into a coherent final clinical state.
The benchmark evaluates more than final-answer classification.
A system must determine:
how each evidence event affects each tracked clinical item;
whether an item should be confirmed… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-dependency-graph-reasoning-v0.1.clinical-intervention-sequencing-and-state-control-v0.2
Clinical Multi-Evidence State Integration Benchmark
CMESI v0.2
The Clinical Multi-Evidence State Integration Benchmark (CMESI) evaluates whether an AI system can reconstruct the evolving state of a complex clinical case across a sequence of heterogeneous evidence events.
CMESI does not test whether a model can identify a diagnosis from a static vignette alone. It tests whether the model can:
maintain several competing clinical hypotheses simultaneously;… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-intervention-sequencing-and-state-control-v0.2.suayptalha__Clarus-7B-v0.3-details
Dataset Card for Evaluation run of suayptalha/Clarus-7B-v0.3
Dataset automatically created during the evaluation run of model suayptalha/Clarus-7B-v0.3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/suayptalha__Clarus-7B-v0.3-details.clinical-constraint-intervention-planning-v0.1Clinical Constraint Intervention Planning v0.1
Clinical Constraint Intervention Planning (CCIP) is a synthetic benchmark for evaluating whether an AI system can construct safe, temporally valid intervention plans under interacting clinical constraints.
The benchmark tests more than selection of a plausible intervention. A model must preserve active constraints, satisfy intervention preconditions at the moment of execution, distinguish treatment initiation from confirmed establishment, respect… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-constraint-intervention-planning-v0.1.suayptalha__Clarus-7B-v0.1-details
Dataset Card for Evaluation run of suayptalha/Clarus-7B-v0.1
Dataset automatically created during the evaluation run of model suayptalha/Clarus-7B-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/suayptalha__Clarus-7B-v0.1-details.suayptalha__Clarus-7B-v0.2-details
Dataset Card for Evaluation run of suayptalha/Clarus-7B-v0.2
Dataset automatically created during the evaluation run of model suayptalha/Clarus-7B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/suayptalha__Clarus-7B-v0.2-details.justina_clarusEnglish
Summary. train17.jsonl contains Portuguese legal Q&A samples. It represents 5% of the data used to train the Clarus model by JUSTINA. Focus areas: Portuguese Civil Procedure Code (CPC), Civil Code, and an intensive subset on abuse of rights.
Languages. pt-PT.
Format. JSON Lines. Each line is one object with a messages array of chat turns:
role: "user" or "assistant".
content: plain text in Portuguese.
No headers, no trailing commas.
Schema.
{
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/VirtuoTuring/justina_clarus.justina_clarus_clean_small
English
Summary
dataset_legal_PT-PT.jsonl is the deduplicated companion of the raw set. Portuguese legal Q&A in chat format. About 95% juridical content across Civil Code, Civil Procedure, corporate, and family law, plus doctrinal discussion. Exact duplicate lines removed to improve signal-to-noise for supervised fine-tuning and evaluation.
Languages
Portuguese (pt-PT)
Format
JSON Lines (.jsonl). Each line is one chat sample with a messages array of… See the full description on the dataset page: https://huggingface.co/datasets/VirtuoTuring/justina_clarus_clean_small.
