datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.BanglaContextualBias
Dataset Card for Bangla Contextual Bias
The Bangla Contextual Bias dataset corresponds to the data described in the paper "An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla" accepted in ACL 2024 (Findings).
Dataset Description
The dataset has different parts for different bias detection experiments conducted for Bengali.
WEAT & SEAT
For the WEAT experiment, the dataset is translated from its English counterpart and… See the full description on the dataset page: https://huggingface.co/datasets/csebuetnlp/BanglaContextualBias.BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.BanglaBook
BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews
This repository contains the code, data, and models of the paper titled "BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews" published in the Findings of the Association for Computational Linguistics: ACL 2023.
License: Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International
Data Format
Each row consists of a book review sample. The… See the full description on the dataset page: https://huggingface.co/datasets/Starscream-11813/BanglaBook.BanglaMultiHate
BanglaMultiHate: Multi-task Bangla Hate-speech Dataset
The BanglaMultiHate dataset collected public comments from YouTube videos using the YouTube API, primarily from Somoy TV, which is a popular Bangla News channel. The comments belong to 19 different categories, including Business, Celebrities, Disaster, Entertainment, Fashion, Geopolitics, Health, History, International, Lifestyle, Literature, Miscellaneous, National, Opinion, Politics, Religion, Science, Sports, and Technology… See the full description on the dataset page: https://huggingface.co/datasets/aridhasan/BanglaMultiHate.bangla-nlp-catalog
Bangla NLP Catalog
A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link.
This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them.
Why this exists
Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.bangla-dialect-normalization
Bangla Dialect Normalization Dataset
A parallel corpus mapping standard Bangla to five regional Bangla dialects,
built from the Vashantor dataset. Each row contains the same sentence in
standard Bangla and Banglish (romanized), alongside its dialect Bangla and
dialect Banglish equivalent, plus an English gloss.
Regions covered
Barishal, Chittagong, Mymensingh, Noakhali, Sylhet
Schema
Field
Description
standard_bangla
Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.bangla-alpaca
Bangla Alpaca
Bangla Alpaca is a culturally localized Bangla (বাংলা) adaptation of the Stanford Alpaca dataset. Unlike simple translation, this dataset uses native-first localization to produce natural, conversational Bangla instruction-following data for training high-quality LLMs.
📊 Overview
Aspect
Description
Language
Bangla (বাংলা)
Format
Instruction-Input-Output
Samples
~52K
License
Apache 2.0
📁 Dataset Structure
{… See the full description on the dataset page: https://huggingface.co/datasets/abubakar-siddik/bangla-alpaca.hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.BanglaCEH
BanglaCEH: A Benchmark for Culturally Entangled Homograph Disambiguation in Bangla
BanglaCEH is the benchmark released with the paper "When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs."
Many Bangla words are simultaneously a personal name and a culturally loaded common noun. মায়া (Maya) is both a common girl's name and a word for deep affectionate compassion; আরিফ (Arif) is a boy's name and… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-xyz/BanglaCEH.BanglaSTEMbangla-english-banglish-pairs
Bangla-English-Banglish Trilingual Pairs
Overview
This dataset provides contrastive training pairs for fine-tuning trilingual (Bangla / Banglish / English) sentence embedding models (such as BGE-M3). It is designed to impart robustness to Banglish spelling variation.
The dataset is combined from two main sources:
LLM-generated Banglish spelling variants.
The OPUS-100 EN-BN parallel corpus.
Included Files
File
Rows
Size
Description… See the full description on the dataset page: https://huggingface.co/datasets/istiaqfuad/bangla-english-banglish-pairs.banglabridge-instructions
Dataset Card — BanglaBridge Banglish Instruction Set
Summary
An original instruction-tuning dataset for code-mixed / romanized Bengali
("Banglish") — the register 100M+ people actually type online
(e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by
safe, deterministic transformation of our own templates. Nothing is scraped, so the
whole set is free to redistribute on Hugging Face and Kaggle.
This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.BangladeshiVQA
BangladeshiVQA
A culturally grounded Bangla Visual Question Answering benchmark.
BangladeshiVQA is a native Bangla VQA benchmark of 2,068 Bangladeshi images and 7,038
open-ended question–answer pairs, organized into three cognitive levels and seven
image categories. To our knowledge it is the first Bangla VQA dataset with dedicated
in-image Bangla scene-text (OCR) questions, and the first to split strictly by image ID to
prevent train/test leakage.
This Hugging Face repository… See the full description on the dataset page: https://huggingface.co/datasets/tanim494/BangladeshiVQA.hsc-biology-bangla-dataset
🌿 HSC Biology Bangla Dataset (Plant Physiology)
The Ultimate Resource for Bengali STEM NLP
This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants.
✨ Key Highlights
Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.bangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.MBPP-Bangla
🐯 MBPP-Bangla: A Benchmark for Evaluating Bangla Code Generation
Accepted at LREC 2026
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri
George Mason University, Fairfax, VA, USA
The first expert-validated, multi-language Bangla code generation benchmark with 974 problems across 5 programming languages.
⚠️ Note: The benchmark will be released after the LREC 2026 conference. Stay tuned!
Overview
MBPP-Bangla is a… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/MBPP-Bangla.alpaca_banglabangla-wikipedia
Bangla (Bengali) Wikipedia Articles Dataset
Request More ScrapesOrder Private Scrapes
Current Progress: Approx 20%
Dataset Summary
This dataset contains a comprehensive extraction of articles from the Bangla (Bengali) Wikipedia. It is designed for Natural Language Processing (NLP) tasks, linguistic research, and training Large Language Models (LLMs) to better understand and generate the Bengali language.
Copyright and Fair Use
I do not own the… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-wikipedia.indian_university_guidance_for_bangladeshi_students
Indian University Guidance for Bangladeshi Students Dataset
Dataset Description
This dataset contains 7,044 high-quality, instruction-formatted Question-Answer pairs designed for fine-tuning Large Language Models (LLMs). The primary goal of this dataset is to create a specialized AI counselor that provides accurate, culturally relevant, and comprehensive guidance on Indian universities for Bangladeshi students.
The dataset was generated through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/millat/indian_university_guidance_for_bangladeshi_students.bangla-llm-data
Bangla NLP Text Corpus — 800K+ Bangla Text Samples for LLM Training and NLP Research
The largest open, multi-domain Bangla text dataset, combining 801,645 samples from 15 different sources — newspapers, social media, education, reviews, QA, medical, poetry, and more. Ready for Bangla LLM pretraining, fine-tuning, and downstream NLP tasks.
Why This Dataset?
Bangla (Bengali) is the 7th most spoken language in the world with 230M+ speakers, but high-quality Bangla NLP… See the full description on the dataset page: https://huggingface.co/datasets/nahidstaq/bangla-llm-data.BanglaMultiHate
BanglaMultiHate: Multi-task Bangla Hate-speech Dataset
The BanglaMultiHate dataset collected public comments from YouTube videos using the YouTube API, primarily from Somoy TV, which is a popular Bangla News channel. The comments belong to 19 different categories, including Business, Celebrities, Disaster, Entertainment, Fashion, Geopolitics, Health, History, International, Lifestyle, Literature, Miscellaneous, National, Opinion, Politics, Religion, Science, Sports, and… See the full description on the dataset page: https://huggingface.co/datasets/Roy2022331060/BanglaMultiHate.Bangla_question_answer_pair_70K_datasetBanglaQuranPunctuationDataset
Bangla Quran Punctuation Dataset
A high-quality dataset for Bangla punctuation restoration, derived exclusively from the Bangla translation of the Holy Quran.
Dataset Description
This dataset is designed for training models on punctuation restoration in Bangla text. Each sample consists of:
human: Bangla text with all punctuation removed
gpt: The original Bangla text with correct punctuation (। , ; : - ! ?)
All samples are derived from the Bangla translation of the… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaQuranPunctuationDataset.Bangla_jokes
Dataset Card for Bangla Jokes Dataset
The Bangla Jokes Dataset is a collection of humorous text samples written in Bengali (Bangla). This dataset is intended for NLP research and model training, especially in the area of Bangla-language humor generation, sentiment, or cultural studies. It is one of the first attempts to gather a sizable dataset of jokes in Bangla for open-source use.
Dataset Details
Curated by: Adnan1837
Funded by: No one
Shared by: Adnan, Md.… See the full description on the dataset page: https://huggingface.co/datasets/adnan1837/Bangla_jokes.bangla-bcs-qsGPTeacher-BanglaBanglaPunctDataset
Bangla Punctuation Restoration Dataset
A merged, high-quality Bangla dataset for punctuation restoration, formatted as instruction-tuning conversation pairs.The dataset is suitable for fine-tuning Large Language Models (LLMs) and sequence models to restore punctuation in Bangla text.
Dataset Summary
Language: Bengali (Bangla)
Task: Punctuation Restoration
Format: JSONL (instruction-style conversations)
Max chunk length: ~256 characters
Punctuation covered:। ! ? , ; : -… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaPunctDataset.xlsum_bangla
Dataset Card for kawsarahmd/papers_summary_datasets_xsum_bangla
This dataset is derived from csebuetnlp/xlsum (subset: bengali).
Dataset Description
Overview
Original dataset: csebuetnlp/xlsum
Subset: bengali
Total samples: 10126
Split source: Original splits from dataset
Splits
Train split: 8102 samples (80.0%)
Validation split: 1012 samples (10.0%)
Test split: 1012 samples (10.0%)
Features
{
"id": "object",
"url": "object"… See the full description on the dataset page: https://huggingface.co/datasets/kawsarahmd/xlsum_bangla.
