mashey/dv-synthetic-errors-lg
Dhivehi Correction Dataset Dataset Description This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. Dataset Summary The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of: A correct Dhivehi sentence The same sentence with synthetic errors Data Splits The dataset is split into: Train: 80% (~5.8M… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dv-synthetic-errors-lg.
04
Duplicate from alakxender/dv-synthetic-errors-lg
