CoolFace
Datasetpublic

alakxender/dv-synthetic-errors-lg

Dhivehi Correction Dataset Dataset Description This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. Dataset Summary The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of: A correct Dhivehi sentence The same sentence with synthetic errors Data Splits The dataset is split into: Train: 80% (~5.8M… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors-lg.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes26downloads
Dataset Card

Dhivehi Correction Dataset

Dataset Description

This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts.

Dataset Summary

The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of:

  • —A correct Dhivehi sentence
  • —The same sentence with synthetic errors

Data Splits

The dataset is split into:

  • —Train: 80% (~5.8M pairs)
  • —Validation: 10% (~720K pairs)
  • —Test: 10% (~720K pairs)

Dataset Structure

Each split contains the following fields:

  • —original: The original, correct sentence
  • —error: The sentence with synthetic errors
  • —error_types: The types of errors introduced
Annotations

Errors are synthetically generated using rule-based transformations including:

  • —Diacritic errors
  • —Character mistakes
  • —Tense marker errors
  • —Case marking mistakes
  • —Word order issues
  • —Agreement errors
  • —Suffix errors