CoolFace
Datasetpublic

mashey/dv-synthetic-errors-lg

Dhivehi Correction Dataset Dataset Description This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. Dataset Summary The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of: A correct Dhivehi sentence The same sentence with synthetic errors Data Splits The dataset is split into: Train: 80% (~5.8M… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dv-synthetic-errors-lg.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes4downloads
1 commits on main
ad8b6d34mo ago

Duplicate from alakxender/dv-synthetic-errors-lg

mashey, alakxender