Badhon/BanglaPunctDataset
Bangla Punctuation Restoration Dataset A merged, high-quality Bangla dataset for punctuation restoration, formatted as instruction-tuning conversation pairs.The dataset is suitable for fine-tuning Large Language Models (LLMs) and sequence models to restore punctuation in Bangla text. Dataset Summary Language: Bengali (Bangla) Task: Punctuation Restoration Format: JSONL (instruction-style conversations) Max chunk length: ~256 characters Punctuation covered:। ! ?… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaPunctDataset.
017
1---2license: apache-2.03task_categories:4- token-classification5- text-classification6tags:7- code8---9 10# Bangla Punctuation Restoration Dataset11 12A merged, high-quality Bangla dataset for **punctuation restoration**, formatted as **instruction-tuning conversation pairs**. 13The dataset is suitable for fine-tuning Large Language Models (LLMs) and sequence models to restore punctuation in Bangla text.14 15---16 17## Dataset Summary18 19- **Language**: Bengali (Bangla)20- **Task**: Punctuation Restoration21- **Format**: JSONL (instruction-style conversations)22- **Max chunk length**: ~256 characters23- **Punctuation covered**: 24 `। ! ? , ; : -`25 26Each example contains:27- An **unpunctuated Bangla text** as input28- A **punctuated Bangla text** as output29 30---31 32## Data Format33 34Each dataset entry follows the structure below:35 36```json37{38 "conversations": [39 {40 "from": "human",41 "value": "এটি একটি উদাহরণ বাক্য যেখানে কোন বিরামচিহ্ন নেই"42 },43 {44 "from": "gpt",45 "value": "এটি একটি উদাহরণ বাক্য, যেখানে কোনো বিরামচিহ্ন নেই।"46 }47 ],48 "source": "source_name",49 "score": 7.8350}51````52 53### Field Description54 55| Field | Description |56| --------------- | -------------------------------------------------- |57| `conversations` | Instruction-style input–output pair |58| `source` | Origin of the sample |59| `score` | Random float (4.0–10.0), placeholder quality score |60 61---62 63## Data Sources64 65This dataset is created by merging and processing three main sources:66 67### 1. Badhon/BanglaQuranPunctuationDataset (Hugging Face)68 69* Bangla Quran text with accurate punctuation70* High grammatical and punctuation consistency71 72### 2. hishab/hishab-pr-bn-v1 (Hugging Face)73 74* Bangla punctuation restoration dataset75* Diverse sentence structures and punctuation styles76 77### 3. Web-Collected Bangla Text (Custom Scraper)78 79High-quality Bangla paragraphs collected from:80 81* Bengali Wikipedia82* Prothom Alo (`prothomalo.com`)83* BBC Bangla (`bbc.com/bengali`)84* Bangla blogs:85 86 * somewhereinblog.net87 * sachalayatan.com88 * amarblog.com89 90---91 92## Data Processing93 94The following preprocessing steps were applied:95 96### Cleaning97 98* Preserved Bangla punctuation (`। ! ? , ; : -`)99* Removed encoding artifacts and noisy symbols100* Ensured Bangla language dominance101 102### Chunking103 104* Text split into chunks of ≤256 characters105* Sentence-boundary–aware splitting106* Avoided mid-sentence truncation107 108### Quality Filtering109 110* Minimum and maximum length thresholds111* Required presence of natural punctuation112* Removed malformed or low-quality samples113 114### Deduplication115 116* Removed duplicates using punctuated-text matching117* Improved diversity and reduced redundancy118 119---120 121## Intended Uses122 123* Fine-tuning LLMs for Bangla punctuation restoration124* Training instruction-following Bangla NLP models125* Sequence-to-sequence or token-classification approaches126* ASR post-processing pipelines (speech → text → punctuation)