datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-real-word-errors
Synthetic Real-Word Error Datasets
This repository contains synthetic German data for grammatical error detection and correction, with a focus on context-dependent real-word errors.
The repository provides four subsets:
Subset
Description
Examples
mixed_real_word
Mixed real-word errors
99,812
capitalization
Capitalization errors
99,664
case
Case errors
99,706
verb
Verb errors
99,780
Each subset contains both erroneous and correct sentences and can therefore… See the full description on the dataset page: https://huggingface.co/datasets/aurorra/synthetic-real-word-errors.quantum-error-mitigation-and-benchmarking
Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking
A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.omnimcp_type_error_mypy_resolver_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_type_error_mypy_resolver_teaser.arabic-grammar-errorsopenslr-sinhala-synthetic-spell-errors-quarter
Sinhala Dyslexic Spelling Correction Dataset
Dataset Description
This dataset contains Sinhala and code-mixed (Sinhala-English) text pairs for training spelling correction models, specifically designed to address dyslexia-like spelling errors.
Features
dyslexic_sentence: Input text with dyslexia-like spelling errors (string)
correct_sentence: Corrected output text (string)
Dataset Statistics
Split
Samples
Train
37,056
Test
9,265… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/openslr-sinhala-synthetic-spell-errors-quarter.Vietnamese_spelling_error
Vietnamese Spelling Error Dataset
This dataset contains examples of Vietnamese text with spelling errors and their corresponding corrections. It is intended to be used for training and evaluating models in spelling correction tasks, particularly for the Vietnamese language.
Dataset Summary
Name: Vietnamese Spelling Error Dataset
Language: Vietnamese
File Format: [CSV/Parquet/dataset/etc.]
Columns:
text: The corresponding corrected version of the text.
error_text: The… See the full description on the dataset page: https://huggingface.co/datasets/ShynBui/Vietnamese_spelling_error.error-detection-positives
error-detection-positives
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines positives samples from multiple domains.
Domain Breakdown
gsm8k: 50 samples
math: 53 samples
metamathqa: 93 samples
orca_math: 96 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math)
question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-positives.error-detection-negatives
error-detection-negatives
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines negatives samples from multiple domains.
Domain Breakdown
gsm8k: 57 samples
math: 44 samples
metamathqa: 59 samples
orca_math: 54 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math)
question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-negatives.error-detection-positives_perturbed
error-detection-positives_perturbed
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines positives_perturbed samples from multiple domains.
Domain Breakdown
gsm8k: 48 samples
math: 42 samples
metamathqa: 72 samples
orca_math: 85 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-positives_perturbed.ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513
GPT-5.5 Medium Reannotation of Qwen3.5-Positive OCR2 Coding Steps
This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and contains GPT-5.5 medium-reasoning reannotations for the 1,536 Qwen3.5-positive error steps.
Summary
Source dataset: JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513
Source rows: 500 K2-Think Codeforces traces
Source manifest-selected Qwen3.5 labels: 10,000 steps
GPT-5.5 reannotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513.DeepScaleR-Qwen3-1.7B-2k-strategy-error-200
DeepScaleR Qwen3 1.7B 2K strategy errors
This dataset contains 200 distinct questions selected from
zjhhhh/DeepScaleR-Qwen3-1.7B-2k-agreed-regraded-le5-coded
at revision 8b6e0f481bced00132c95fb631745d4992fa19fd.
Each row has one manually selected model response whose main failure is a
strategy error relative to the source row's code_hint: the response does not
materially use the hint's core route, substitutes another strategy, or omits a
decisive hinted stage in favor of an… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/DeepScaleR-Qwen3-1.7B-2k-strategy-error-200.AstralMath-v1-ErrorTracesChangelog:
2026-03-26:
Public AstralMath-v1-ErrorTraces, include 520k error traces that models encounter during the synthesis process.
2026-03-25:
Add new 50k datapoints, replace ~10k old datapoints with higher quality synthetic questions(12 consensus tranform use tool for verify) to stage 1.
Removed ~600 datapoints affect by extract function bug(raise incomplete question).
Replace 1 question in AstralBench(hmmt-feb-2026-algebra-p7 -> open-rl-combinatorics-247678).
2026-03-12:
Release… See the full description on the dataset page: https://huggingface.co/datasets/nguyen599/AstralMath-v1-ErrorTraces.
