CoolFace
Datasetpublic

NLPC-UOM/anonymized-sinhala-letter-corpus

Anonymized Sinhala Official Letter Corpus A small, hand-curated corpus of 151 formal Sinhala letters, fully anonymized with bracketed placeholders. It is intended for training and evaluating models that generate, complete, or classify Sinhala official correspondence — a task with very little public training data. Dataset at a glance Examples 151 Language Sinhala (si) Register Formal throughout Letter length 42–240 words (median 108, mean 114)… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/anonymized-sinhala-letter-corpus.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes8downloads

NLPC-UOM/anonymized-sinhala-letter-corpus · main · files are served by the source, never re-hosted here