CoolFace
Datasetpublic

alakxender/dv-synthetic-errors

DV Text Errors Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools. About Dataset Task: Text error correction Language: Dhivehi (dv) Dataset Structure Input-output pairs of Dhivehi text: correct: Original correct sentences incorrect: Sentences with synthetic errors Statistics Train set: {train_examples}… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes44downloads
Dataset Card

DV Text Errors

Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools.

About Dataset

  • Task: Text error correction
  • Language: Dhivehi (dv)

Dataset Structure

Input-output pairs of Dhivehi text:

  • correct: Original correct sentences
  • incorrect: Sentences with synthetic errors

Statistics

  • Train set: {train_examples} examples ({0.7999997975429817}%)
  • Test set: {test_examples} examples ({0.10000010122850919}%)
  • Validation set: {val_examples} examples ({0.10000010122850919}%)

Details:

  • Unique words: {448628}
json
{
  "total_examples": {
    "train": 3161164,
    "test": 395146,
    "validation": 395146
  },
  "avg_sentence_length": {
    "train": 11.968980097204701,
    "test": 11.961302910822836,
    "validation": 11.973824864733542
  },
  "error_distribution": {
    "min": 0,
    "max": 2411,
    "avg": 64.85144965588626
  }
}

Usage

python
from datasets import load_dataset

dataset = load_dataset("alakxender/dv-synthetic-errors")

Dataset Creation

Created using:

  • Source: Collection of Dhivehi articles
  • Error generation: Character and diacritic substitutions
  • Error rate: 30% per word probability