CoolFace
Datasetpublic

mashey/dv-synthetic-errors-lg

Dhivehi Correction Dataset Dataset Description This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. Dataset Summary The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of: A correct Dhivehi sentence The same sentence with synthetic errors Data Splits The dataset is split into: Train: 80% (~5.8M… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dv-synthetic-errors-lg.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes4downloads

mashey/dv-synthetic-errors-lg · main · files are served by the source, never re-hosted here