datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
titulm-bangla-corpus
TituLM Bangla Corpus
This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability.
This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/hishab/titulm-bangla-corpus.BanglaEng-SynCorpus
BanglaEng-SynCorpus
Dataset Summary
BanglaEng-SynCorpus is a large-scale synthetic Bangla–English parallel corpus designed to support research in Neural Machine Translation (NMT) and other Bangla–English bilingual NLP tasks.The corpus is generated using linguistically validated sentence templates combined with topic-wise curated vocabularies, covering all 12 English/Bangla tense structures.
Due to extreme scale (trillions of possible sentence pairs), the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Eamin-sust/BanglaEng-SynCorpus.titulm-bangla-corpus
TituLM Bangla Corpus
This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability.
This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we… See the full description on the dataset page: https://huggingface.co/datasets/shofikul-1234/titulm-bangla-corpus.bangla-instruction-dataset
🧠 Bangla Instruction Dataset
This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models.
📚 Dataset Splits
The dataset is organized into the following splits:
Split Name
Source Dataset
Description
OdiaGenAI
OdiaGenAI/all_combined_bengali_252k
A large-scale collection of diverse Bangla instructions and responses.
chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.BanglaVerse
Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects
Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision–language… See the full description on the dataset page: https://huggingface.co/datasets/FaiyazAbdullah114708/BanglaVerse.BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.Bangla-TextBook
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
---
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.bangla-corpus
TituLM Bangla Corpus
This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability.
This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/munzurul/bangla-corpus.Bangla-Instruct
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.Institutional-Information-of-Bangladesh
Institutional-Information-of-Bangladesh Dataset
This Dataset contains all verified and authorized Institutional information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.BanglaSEC
BanglaSEC
A 1.18M-pair parallel corpus for Bangla spelling error correction, with character-level error masks across 14 error types.
BanglaSEC is the corpus introduced in A transformer based spelling error correction framework for Bangla and resource scarce Indic languages (Bijoy, Hossain, Islam & Shatabda, Computer Speech & Language 89:101703, 2025). Each row pairs a correct Bangla word with an erroneous form, labelled by error type and annotated with a binary mask marking… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaSEC.bangladesh-scob-judgment-summarization
Bangladesh Supreme Court (SCOB) High Court Division Judgment Summarization Dataset
Dataset Summary
The Bangladesh Supreme Court (SCOB) High Court Division Judgment Summarization Dataset is a curated, high-quality legal NLP dataset comprising all 235 canonical judgments published in the Supreme Court Online Bulletin (SCOB) by the High Court Division of the Supreme Court of Bangladesh.
Each sample pairs a complete, cleaned legal judgment body with its official… See the full description on the dataset page: https://huggingface.co/datasets/Hasin2026/bangladesh-scob-judgment-summarization.BanglaSleep-CoT
BanglaSleep-CoT
The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces.
Built for the Uncharted Data Challenge by Adaption Labs.
Expanded using Adaptive Data by Adaption.
Dataset at a Glance
Why This Dataset Exists
Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.alpaca-gpt4-bangla
alpaca-gpt4-bangla
A Bangla (Bengali) instruction-following dataset for supervised fine-tuning (SFT) of large language models. It contains ~49,969 instruction-response pairs covering a wide range of topics -- coding, creative writing, reasoning, summarization, math, open Q&A -- suitable for teaching a base model to follow instructions in Bangla.
This dataset is a Korean -> Bangla machine translation of FreedomIntelligence/alpaca-gpt4-korean, which is itself a Korean translation… See the full description on the dataset page: https://huggingface.co/datasets/ihumaunkabir/alpaca-gpt4-bangla.BanglaGEC
BanglaGEC: A Large-Scale Parallel Corpus for Bangla Grammatical Error Correction
BanglaGEC is a large-scale parallel corpus of 7,074,425 (~7.1M) sentence pairs for Bangla (Bengali) Grammatical Error Correction (GEC). Each pair maps a grammatically erroneous Bangla sentence to its grammatically correct counterpart, along with the error type, making it directly usable for training and evaluating sequence-to-sequence models, transformers, and large language models on the Bangla… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BanglaGEC.bangla-alpaca
Bangla Alpaca
Bangla Alpaca is a culturally localized Bangla (বাংলা) adaptation of the Stanford Alpaca dataset. Unlike simple translation, this dataset uses native-first localization to produce natural, conversational Bangla instruction-following data for training high-quality LLMs.
📊 Overview
Aspect
Description
Language
Bangla (বাংলা)
Format
Instruction-Input-Output
Samples
~52K
License
Apache 2.0
📁 Dataset Structure
{… See the full description on the dataset page: https://huggingface.co/datasets/abubakar-siddik/bangla-alpaca.BanglaPRCorpus
BanglaPRCorpus
A 1.48M-pair corpus for Bangla punctuation restoration — unpunctuated source sentences paired with their fully punctuated targets, labelled by how many punctuation marks were removed.
BanglaPRCorpus is the corpus introduced in Advancing Bangla Punctuation Restoration by a Monolingual Transformer-Based Method and a Large-Scale Corpus (Bijoy et al., EMNLP 2023 Workshop on Bangla Language Processing), alongside the Jatikarok model.
Each row is a (source, target)… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaPRCorpus.hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.ekpatagolpo-scrape-bangla-literature
Ekpatagolpo Bengali Stories Archive
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a large-scale, curated text dataset scraped from ekpatagolpo.com. The primary goal of this archive is to preserve a massive collection of purely human-written Bengali literature and stories (Bangla Golpo), creating a distinct record of human creativity separate from AI-generated text.
Purpose and Usage
This dataset is published publicly under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/ekpatagolpo-scrape-bangla-literature.BanglaCEH
BanglaCEH: A Benchmark for Culturally Entangled Homograph Disambiguation in Bangla
BanglaCEH is the benchmark released with the paper "When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs."
Many Bangla words are simultaneously a personal name and a culturally loaded common noun. মায়া (Maya) is both a common girl's name and a word for deep affectionate compassion; আরিফ (Arif) is a boy's name and… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-xyz/BanglaCEH.banglabridge-instructions
Dataset Card — BanglaBridge Banglish Instruction Set
Summary
An original instruction-tuning dataset for code-mixed / romanized Bengali
("Banglish") — the register 100M+ people actually type online
(e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by
safe, deterministic transformation of our own templates. Nothing is scraped, so the
whole set is free to redistribute on Hugging Face and Kaggle.
This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.jugantor.com-scrape-bangla
Jugantor News Archive (Bangla)
Overview
This repository contains a comprehensive text dataset scraped from jugantor.com, one of the leading Bengali daily newspapers in Bangladesh. The primary goal of this archive is to preserve a massive collection of purely human-written journalism, editorials, and news reports, creating a distinct record of human-authored text separate from AI-generated content.
Purpose and Usage
This dataset is published publicly and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/jugantor.com-scrape-bangla.DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-DataDeepSeek R1 Bangla MMLU Distil Dataset
Original Dataset: hishab/bangla-mmlu
Train Samples: 17,796
Test Samples: 2,576
Total API Cost: 7K BDT
Contributors:
Myself
Numaer
How the Dataset was created
Step 1 - Base Dataset
I've used bangla-mmlu dataset released by hisab. Kudos to them for creating and open sourcing the dataset. Without their dataset this synthetic reasoning dataset won't exist in the first place.
Step 2 - Select Subset
Since I'm… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-Data.hsc-biology-bangla-dataset
🌿 HSC Biology Bangla Dataset (Plant Physiology)
The Ultimate Resource for Bengali STEM NLP
This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants.
✨ Key Highlights
Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.bangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.MBPP-Bangla
🐯 MBPP-Bangla: A Benchmark for Evaluating Bangla Code Generation
Accepted at LREC 2026
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri
George Mason University, Fairfax, VA, USA
The first expert-validated, multi-language Bangla code generation benchmark with 974 problems across 5 programming languages.
⚠️ Note: The benchmark will be released after the LREC 2026 conference. Stay tuned!
Overview
MBPP-Bangla is a… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/MBPP-Bangla.bangla-health-related-paraphrased-dataset
Dataset Card for "BanglaHealthParaphrase"
BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.Bangla-SFT-50k
Bangla-SFT
Bangla-SFT is an instruction-following dataset containing 50,053 Bengali prompt-response pairs. It was scaled up from a 500-sample seed dataset (spitfire4794/bang_seed).
Dataset Summary
The dataset covers 6 task categories. The prompts are designed to be self-contained (hydrated with appropriate contextual inputs), and the responses are formatted to be direct, omitting conversational prefaces and filler.
Seed Generation: Baseline instructions generated… See the full description on the dataset page: https://huggingface.co/datasets/spitfire4794/Bangla-SFT-50k.
