CoolFace
Datasetpublic

mehedihasanbijoy/BanglaPRCorpus

BanglaPRCorpus A 1.48M-pair corpus for Bangla punctuation restoration — unpunctuated source sentences paired with their fully punctuated targets, labelled by how many punctuation marks were removed. BanglaPRCorpus is the corpus introduced in Advancing Bangla Punctuation Restoration by a Monolingual Transformer-Based Method and a Large-Scale Corpus (Bijoy et al., EMNLP 2023 Workshop on Bangla Language Processing), alongside the Jatikarok model. Each row is a (source, target)… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaPRCorpus.

sourceHugging Facemitupdated 15d agoView on Hugging Face
0likes60downloads

mehedihasanbijoy/BanglaPRCorpus · main · files are served by the source, never re-hosted here