CoolFace
Datasetpublic

nrl-ai/vn-spell-correction-eval-real

vn-spell-correction-eval-real Out-of-distribution evaluation corpus for Vietnamese spell-correction models — 150 hand-curated (noisy, clean) pairs sampled from real VN error sources, not generated by nom.text.noise. This is the test set we use to verify a spell-correction model generalises beyond its own synthetic training distribution. A model that scores 95 % on nom-vn's synthetic eval grid and 60 % on this set is overfit to the noise generator. Splits… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
0likes126downloads
settings

This repository belongs to nrl-ai on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namevn-spell-correction-eval-real
visibilitypublic
licencecc0-1.0
gatedno
ownernrl-ai
Account settings
nrl-ai/vn-spell-correction-eval-real · CoolFace