Badhon/BanglaPunctDataset
Bangla Punctuation Restoration Dataset A merged, high-quality Bangla dataset for punctuation restoration, formatted as instruction-tuning conversation pairs.The dataset is suitable for fine-tuning Large Language Models (LLMs) and sequence models to restore punctuation in Bangla text. Dataset Summary Language: Bengali (Bangla) Task: Punctuation Restoration Format: JSONL (instruction-style conversations) Max chunk length: ~256 characters Punctuation covered:। ! ?… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaPunctDataset.
017
