boundary-integrity
category-boundary-integrity-v01Cardinal Meta Dataset 3.2Category Boundary Integrity
Purpose
Test whether claims stay within the correct category
Test whether distinct categories are not collapsed
Test whether certainty is not imported across category boundaries
Central question
What category is this claim in
What this dataset catches
Legal treated as ethical
Statistical treated as causal
Performance treated as safety
Procedure treated as morality
Reporting treated as impact
Consistency treated as truth… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/category-boundary-integrity-v01.tts-semantic-boundary-integrity-v0.1
What this dataset tests
Speech must preserve boundaries.
Negation matters.
Modality matters.
Conditions matter.
Numbers matter.
Why it exists
Voice systems can blur meaning.
May becomes will.
If disappears.
Only gets lost.
Numbers get rounded.
This set detects boundary loss.
Data format
Each row contains
source_text
boundary_markers
tts_transcript_with_marks
boundary_pressure
Inline marks stand in for audible emphasis.
What is scored… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/tts-semantic-boundary-integrity-v0.1.boundary-scope-integrity-v01Cardinal Meta Dataset Set 2Boundary and Scope Integrity
Purpose
Test whether models respect evidential limits
Test whether models refuse to answer outside scope
Test whether models separate evidence from inference
Central question
Is this claim inside what can be supported from the given frame
Why this set exists
Assumptions can be named yet still overreach
Reasoning can be valid but applied outside bounds
Scope discipline is a distinct failure mode
What this dataset catches… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/boundary-scope-integrity-v01.population-boundary-integrity-v0.1What this dataset tests
Studied cohort versus claimed population.
It detects generalization creep.
It flags inference beyond enrollment criteria.
Required outputs
cohort_profile
claimed_population_profile
boundary_violations
required_population_disambiguation
risk_if_overgeneralized
Typical failures
assuming age or comorbidity transfer
extending findings to excluded groups
conflating inpatient evidence with outpatient claims
Suggested prompt wrapper
System
You check… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/population-boundary-integrity-v0.1.assumption-boundary-integrity-v0.1What this is
This dataset tests whether a model keeps causal reasoning clean.
If the intervention did not occur,the model must not claim it did.
Unexpected outcomes do not license invented causes.
What it measures
• whether the model confuses correlation with intervention• whether it invents hidden actions to keep a story tidy• whether it can hold uncertainty without collapsing into a false cause
Expected model output
Return only:
AorB
Scoring
Use the shared Clarus scorer.
Metrics:
• accuracy•… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/assumption-boundary-integrity-v0.1.clinical-population-boundary-integrity-v0.1What this dataset tests
Whether clinical claims stay withinthe population actually studied.
Required outputs
studied population
claimed population
boundary violations
safety risk of overreach
Typical failures
adult data applied to children
fit trial populations applied to frail patients
disease subtype expansion without evidence
Suggested prompt wrapper
System
You enforce population boundaries.
You prevent unsafe generalization.
User
Studied… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-population-boundary-integrity-v0.1.
