datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
99-GEO-Errors
99 Errors in GEO
Why organisations become invisible, misrepresented or unsupported in AI answers
GEO means Generative Engine Optimization. This six-language companion book turns 99 recurring representation failures into auditable warnings. Each warning records the evidence needed, a correction protocol, a revalidation question and a machine-readable rule.
Start reading: Open the English PDF · Choose one of six languages · Cite the DOI
Kaan Muraz · NobleJackal ·… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/99-GEO-Errors.toulmin_errors
Reasoning Rubrics — Toulmin-Typed Error Localization Benchmark
A multi-domain benchmark for studying typed reasoning errors in LLMs and AI
scientific reasoning agents. Errors are labeled along four Toulmin
argumentation dimensions: Grounds (premises/facts), Warrant
(inferential step), Qualifier (scope/certainty), Rebuttal
(competing evidence).
The benchmark has two parts:
Typed external benchmarks. Existing reasoning-error benchmarks
relabeled with Toulmin dimensions on top of the… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/toulmin_errors.ct-dosing-errors-benchmarkct-dosing-errors
ds4dh/ct-dosing-errors
Version
This repository contains dataset version 0.2.3.
License
This dataset is licensed under CC BY 4.0 (cc-by-4.0).
toulmin_errors
Toulmin-Errors: A Benchmark for Typed Reasoning-Error Detection
Reasoning-error benchmarks mostly measure factual and logical mistakes.
They rarely measure two argument-level failures: getting the scope of a
claim wrong, and ignoring counter-evidence. In Toulmin's argument model
these are Qualifier (Q) and Rebuttal (R) failures. This benchmark
provides the data to study them, with every error typed along four Toulmin
dimensions: Grounds (premises/facts), Warrant (inferential step)… See the full description on the dataset page: https://huggingface.co/datasets/anonupload1ng/toulmin_errors.qwen3_0.6b-task738_augmented_finding_errors_in_reasoning_traces_Mar16-1513_blendedErrorSpanAnnotation-for-Taiwanese-Hokkien
Error Span Annotation for Taiwanese Hokkien
The Taiwanese Hokkien subset of the SiniticMTError benchmark (Liu et al., 2026).
Human-annotated machine-translation error-span evaluation data for the
Mandarin → Taiwanese Hokkien (Tâi-gí) direction. Each instance contains a
Mandarin source sentence, a Taiwanese Hokkien machine translation, a reference
translation, and expert error-span annotations with severity labels and a
segment-level quality score.
Language pair: Mandarin (zh) →… See the full description on the dataset page: https://huggingface.co/datasets/350016z/ErrorSpanAnnotation-for-Taiwanese-Hokkien.tulu_3_all_errors_updfinepdfs-cy-errors
Dataset Card: finepdfs-cy-errors
Description
This dataset contains Welsh-language text extracted from PDFs using rolmOCR, with automated spelling and grammar error annotations generated by Cysill (the Welsh spell checker). The dataset is derived from the Welsh (cym_Latn) subset of HuggingFaceFW/finepdfs, filtered to include only documents processed with the rolmOCR extractor.
Dataset Statistics
Corpus Size
Total number of words: 1,575,836
Total… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/finepdfs-cy-errors.qwen3_0.6b_finding_errors_in_reasoning_traces_Mar17-1951_blendedqwen3_0.6b-task738_augmented_finding_errors_in_reasoning_traces_Mar16-1513qwen3_0.6b_finding_errors_in_reasoning_traces_Mar17-1951GSM-Plus_model_generations_errorskodcode-complete_1000_qwen7b_att_iter0_att40_sol5_relabeled_dedup_assertion_errorsGSM-Symbolic_model_generations_errorsagreement-errors-ro-rrtagreement-errors-model-resultssynth_citation_errorstranscription-errors-correction
Transcription Errors Correction
Feature Distribution
