CoolFace
Datasetpublic

nrl-ai/vn-spell-correction-eval-real

vn-spell-correction-eval-real Out-of-distribution evaluation corpus for Vietnamese spell-correction models — 150 hand-curated (noisy, clean) pairs sampled from real VN error sources, not generated by nom.text.noise. This is the test set we use to verify a spell-correction model generalises beyond its own synthetic training distribution. A model that scores 95 % on nom-vn's synthetic eval grid and 60 % on this set is overfit to the noise generator. Splits… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
0likes126downloads
11 commits on main
cb763135mo ago

Upload README.md with huggingface_hub

vietanhdev
21564a45mo ago

Upload README.vi.md with huggingface_hub

vietanhdev
c9d56f25mo ago

Upload README.md with huggingface_hub

vietanhdev
39720185mo ago

Publish OOD spell-correction eval (150 sentences, 6 registers)

vietanhdev
546c7075mo ago

Upload telex_real_25.jsonl with huggingface_hub

vietanhdev
34a3ece5mo ago

Upload ocr_25.jsonl with huggingface_hub

vietanhdev
82148845mo ago

Upload news_real_25.jsonl with huggingface_hub

vietanhdev
353b9f25mo ago

Upload mobile_25.jsonl with huggingface_hub

vietanhdev
ed5d9f95mo ago

Upload legal_real_25.jsonl with huggingface_hub

vietanhdev
d246c085mo ago

Upload forum_25.jsonl with huggingface_hub

vietanhdev
0ab6cfb5mo ago

initial commit

vietanhdev