toulmin
Datasets
All datasets matching “toulmin”toulmin_errors
Reasoning Rubrics — Toulmin-Typed Error Localization Benchmark
A multi-domain benchmark for studying typed reasoning errors in LLMs and AI
scientific reasoning agents. Errors are labeled along four Toulmin
argumentation dimensions: Grounds (premises/facts), Warrant
(inferential step), Qualifier (scope/certainty), Rebuttal
(competing evidence).
The benchmark has two parts:
Typed external benchmarks. Existing reasoning-error benchmarks
relabeled with Toulmin dimensions on top of the… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/toulmin_errors.toulmin_errors
Toulmin-Errors: A Benchmark for Typed Reasoning-Error Detection
Reasoning-error benchmarks mostly measure factual and logical mistakes.
They rarely measure two argument-level failures: getting the scope of a
claim wrong, and ignoring counter-evidence. In Toulmin's argument model
these are Qualifier (Q) and Rebuttal (R) failures. This benchmark
provides the data to study them, with every error typed along four Toulmin
dimensions: Grounds (premises/facts), Warrant (inferential step)… See the full description on the dataset page: https://huggingface.co/datasets/anonupload1ng/toulmin_errors.
