CoolFace
Datasetpublic

mehedihasanbijoy/BanglaPRCorpus

BanglaPRCorpus A 1.48M-pair corpus for Bangla punctuation restoration — unpunctuated source sentences paired with their fully punctuated targets, labelled by how many punctuation marks were removed. BanglaPRCorpus is the corpus introduced in Advancing Bangla Punctuation Restoration by a Monolingual Transformer-Based Method and a Large-Scale Corpus (Bijoy et al., EMNLP 2023 Workshop on Bangla Language Processing), alongside the Jatikarok model. Each row is a (source, target)… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaPRCorpus.

sourceHugging Facemitupdated 15d agoView on Hugging Face
0likes60downloads
3 commits on main
113ab9415d ago

Update README.md

mehedihasanbijoy
05032e915d ago

Add files using upload-large-folder tool

mehedihasanbijoy
b620ad915d ago

initial commit

mehedihasanbijoy