CoolFace
Datasetpublic

Badhon/BanglaPunctDataset

Bangla Punctuation Restoration Dataset A merged, high-quality Bangla dataset for punctuation restoration, formatted as instruction-tuning conversation pairs.The dataset is suitable for fine-tuning Large Language Models (LLMs) and sequence models to restore punctuation in Bangla text. Dataset Summary Language: Bengali (Bangla) Task: Punctuation Restoration Format: JSONL (instruction-style conversations) Max chunk length: ~256 characters Punctuation covered:। ! ?… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaPunctDataset.

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
0likes17downloads

Badhon/BanglaPunctDataset · main · files are served by the source, never re-hosted here