alakxender/dv-synthetic-errors-lg
Dhivehi Correction Dataset Dataset Description This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. Dataset Summary The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of: A correct Dhivehi sentence The same sentence with synthetic errors Data Splits The dataset is split into: Train: 80% (~5.8M… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors-lg.
Dhivehi Correction Dataset
Dataset Description
This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts.
Dataset Summary
The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of:
- A correct Dhivehi sentence
- The same sentence with synthetic errors
Data Splits
The dataset is split into:
- Train: 80% (~5.8M pairs)
- Validation: 10% (~720K pairs)
- Test: 10% (~720K pairs)
Dataset Structure
Each split contains the following fields:
original: The original, correct sentenceerror: The sentence with synthetic errorserror_types: The types of errors introduced
Annotations
Errors are synthetically generated using rule-based transformations including:
- Diacritic errors
- Character mistakes
- Tense marker errors
- Case marking mistakes
- Word order issues
- Agreement errors
- Suffix errors
