datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.bangla-alpaca
Bangla Alpaca
Bangla Alpaca is a culturally localized Bangla (বাংলা) adaptation of the Stanford Alpaca dataset. Unlike simple translation, this dataset uses native-first localization to produce natural, conversational Bangla instruction-following data for training high-quality LLMs.
📊 Overview
Aspect
Description
Language
Bangla (বাংলা)
Format
Instruction-Input-Output
Samples
~52K
License
Apache 2.0
📁 Dataset Structure
{… See the full description on the dataset page: https://huggingface.co/datasets/abubakar-siddik/bangla-alpaca.hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.BanglaCEH
BanglaCEH: A Benchmark for Culturally Entangled Homograph Disambiguation in Bangla
BanglaCEH is the benchmark released with the paper "When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs."
Many Bangla words are simultaneously a personal name and a culturally loaded common noun. মায়া (Maya) is both a common girl's name and a word for deep affectionate compassion; আরিফ (Arif) is a boy's name and… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-xyz/BanglaCEH.banglabridge-instructions
Dataset Card — BanglaBridge Banglish Instruction Set
Summary
An original instruction-tuning dataset for code-mixed / romanized Bengali
("Banglish") — the register 100M+ people actually type online
(e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by
safe, deterministic transformation of our own templates. Nothing is scraped, so the
whole set is free to redistribute on Hugging Face and Kaggle.
This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.hsc-biology-bangla-dataset
🌿 HSC Biology Bangla Dataset (Plant Physiology)
The Ultimate Resource for Bengali STEM NLP
This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants.
✨ Key Highlights
Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.bangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.MBPP-Bangla
🐯 MBPP-Bangla: A Benchmark for Evaluating Bangla Code Generation
Accepted at LREC 2026
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri
George Mason University, Fairfax, VA, USA
The first expert-validated, multi-language Bangla code generation benchmark with 974 problems across 5 programming languages.
⚠️ Note: The benchmark will be released after the LREC 2026 conference. Stay tuned!
Overview
MBPP-Bangla is a… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/MBPP-Bangla.bangla-wikipedia
Bangla (Bengali) Wikipedia Articles Dataset
Request More ScrapesOrder Private Scrapes
Current Progress: Approx 20%
Dataset Summary
This dataset contains a comprehensive extraction of articles from the Bangla (Bengali) Wikipedia. It is designed for Natural Language Processing (NLP) tasks, linguistic research, and training Large Language Models (LLMs) to better understand and generate the Bengali language.
Copyright and Fair Use
I do not own the… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-wikipedia.bangla-stories-dataset
Bangla Multi-Task Benchmark Dataset
A comprehensive multi-task dataset for evaluating and fine-tuning LLMs on Bangla.
Statistics
Metric
Value
Total Samples
1,592
Task Types
8
Training
1,273
Validation
159
Test
160
Task Types
Task
Count
Difficulty
text_generation
200
Medium
title_generation
200
Easy
summarization
200
Medium
comprehension
200
Hard
classification
200
Easy
sentiment
200
Easy
keywords
200
Easy… See the full description on the dataset page: https://huggingface.co/datasets/likhonhfai/bangla-stories-dataset.bangla-kobita-scrape-bangla-literature
Bangla Kobita Poetry Archive
Overview
This repository contains a curated text dataset of Bengali poetry scraped from the web, primarily targeting comprehensive poetry platforms like bangla-kobita.com. The primary goal of this archive is to preserve a rich collection of purely human-written Bengali poems (Bangla Kobita), creating a distinct record of human artistic expression, emotion, and linguistic rhythm separate from AI-generated text.
Purpose and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-kobita-scrape-bangla-literature.
