CoolFace
Datasetpublic

mehedihasanbijoy/BanglaSEC

BanglaSEC A 1.18M-pair parallel corpus for Bangla spelling error correction, with character-level error masks across 14 error types. BanglaSEC is the corpus introduced in A transformer based spelling error correction framework for Bangla and resource scarce Indic languages (Bijoy, Hossain, Islam & Shatabda, Computer Speech & Language 89:101703, 2025). Each row pairs a correct Bangla word with an erroneous form, labelled by error type and annotated with a binary mask marking… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaSEC.

sourceHugging Facemitupdated 15d agoView on Hugging Face
0likes112downloads
4 commits on main
78e287a15d ago

Update README.md

mehedihasanbijoy
752dc1115d ago

Update README.md

mehedihasanbijoy
c0b8fbd15d ago

Add BanglaSEC train/valid/test splits (DPCSpell protocol)

mehedihasanbijoy
024253a15d ago

initial commit

mehedihasanbijoy