datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
99-GEO-Errors
99 Errors in GEO
Why organisations become invisible, misrepresented or unsupported in AI answers
GEO means Generative Engine Optimization. This six-language companion book turns 99 recurring representation failures into auditable warnings. Each warning records the evidence needed, a correction protocol, a revalidation question and a machine-readable rule.
Start reading: Open the English PDF · Choose one of six languages · Cite the DOI
Kaan Muraz · NobleJackal ·… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/99-GEO-Errors.toulmin_errors
Reasoning Rubrics — Toulmin-Typed Error Localization Benchmark
A multi-domain benchmark for studying typed reasoning errors in LLMs and AI
scientific reasoning agents. Errors are labeled along four Toulmin
argumentation dimensions: Grounds (premises/facts), Warrant
(inferential step), Qualifier (scope/certainty), Rebuttal
(competing evidence).
The benchmark has two parts:
Typed external benchmarks. Existing reasoning-error benchmarks
relabeled with Toulmin dimensions on top of the… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/toulmin_errors.toulmin_errors
Toulmin-Errors: A Benchmark for Typed Reasoning-Error Detection
Reasoning-error benchmarks mostly measure factual and logical mistakes.
They rarely measure two argument-level failures: getting the scope of a
claim wrong, and ignoring counter-evidence. In Toulmin's argument model
these are Qualifier (Q) and Rebuttal (R) failures. This benchmark
provides the data to study them, with every error typed along four Toulmin
dimensions: Grounds (premises/facts), Warrant (inferential step)… See the full description on the dataset page: https://huggingface.co/datasets/anonupload1ng/toulmin_errors.ErrorSpanAnnotation-for-Taiwanese-Hokkien
Error Span Annotation for Taiwanese Hokkien
The Taiwanese Hokkien subset of the SiniticMTError benchmark (Liu et al., 2026).
Human-annotated machine-translation error-span evaluation data for the
Mandarin → Taiwanese Hokkien (Tâi-gí) direction. Each instance contains a
Mandarin source sentence, a Taiwanese Hokkien machine translation, a reference
translation, and expert error-span annotations with severity labels and a
segment-level quality score.
Language pair: Mandarin (zh) →… See the full description on the dataset page: https://huggingface.co/datasets/350016z/ErrorSpanAnnotation-for-Taiwanese-Hokkien.synth_citation_errors
