CoolFace
Datasetpublic

Badhon/BanglaPunctDataset

Bangla Punctuation Restoration Dataset A merged, high-quality Bangla dataset for punctuation restoration, formatted as instruction-tuning conversation pairs.The dataset is suitable for fine-tuning Large Language Models (LLMs) and sequence models to restore punctuation in Bangla text. Dataset Summary Language: Bengali (Bangla) Task: Punctuation Restoration Format: JSONL (instruction-style conversations) Max chunk length: ~256 characters Punctuation covered:। ! ?… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaPunctDataset.

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
0likes17downloads
3 commits on main
f2fc2269mo ago

Upload BanglaPunctDataset.jsonl

Badhon
12134289mo ago

Update README.md

Badhon
75c560d9mo ago

initial commit

Badhon