datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangla-crime-investigation-patterns-v2
Bangla Crime Investigation Patterns V2
This dataset is an anonymized, structured extraction from 1,078 public Bangla crime-investigation video transcripts. It was built to support machine-learning and LLM research on crime-pattern information extraction, case summarization, event sequencing, entity/relationship extraction, and Bengali investigative-report analysis.
The public release does not include raw full transcripts. It contains short redacted evidence snippets linked to… See the full description on the dataset page: https://huggingface.co/datasets/tanziro/bangla-crime-investigation-patterns-v2.BanglaRQABanglaRQA is a human-annotated Bangla Question Answering (QA) dataset with diverse question-answer types.bangla-instruction-dataset
🧠 Bangla Instruction Dataset
This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models.
📚 Dataset Splits
The dataset is organized into the following splits:
Split Name
Source Dataset
Description
OdiaGenAI
OdiaGenAI/all_combined_bengali_252k
A large-scale collection of diverse Bangla instructions and responses.
chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.BanglaVerse
Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects
Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision–language… See the full description on the dataset page: https://huggingface.co/datasets/FaiyazAbdullah114708/BanglaVerse.bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.Institutional-Information-of-Bangladesh
Institutional-Information-of-Bangladesh Dataset
This Dataset contains all verified and authorized Institutional information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.bangla-math-chat
bangla-math-chat
A math dataset for fine-tuning LLMs to chat on math problems in Bangla. This dataset is a reformatted version of
BanglaLLM/bangla_math_by_Ashrafur.
The code to reformat the original dataset can be found on Github: ShawonAshraf/bangla-math-chat
bangla-mmlu
Data Summary
We curated multiple-choice questions from various open-source educational websites and textbooks, inspired by the original MMLU dataset (Hendrycks et al., 2020). The dataset includes multiple-choice questions from different Bangladeshi exams, such as job exams, the Bangladesh Civil Service Exam, and undergraduate admission exams. In Figure 7, we report category wise distributions.
Bangla MMLU Dataset Overview
The Bangla MMLU dataset consists of a total of 116,503… See the full description on the dataset page: https://huggingface.co/datasets/hishab/bangla-mmlu.bangla-nlp-catalog
Bangla NLP Catalog
A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link.
This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them.
Why this exists
Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.BanglaSocialBias
Dataset Card for Bangla Contextual Bias
The Bangla Social Bias dataset comprises of the data used in the paper titled "Social Bias in Large Language Models For Bangla: An Empirical Study on Gender and Religious Bias".
Dataset Description
The dataset contains different domains of data used for the experimentations mentioned in the paper. A summary of the different categories of data provided in this dataset are:
the formatted raw data collected from open source for the… See the full description on the dataset page: https://huggingface.co/datasets/csebuetnlp/BanglaSocialBias.index-of-the-bangladesh-code
Index of the Bangladesh Code
A structured, machine-readable research index of Bangladesh Code legal records.
Maintainer: Afzal Hosen MandalOrganization: LegalDefenseHub / Afzal & AssociatesJurisdiction: BangladeshCoverage: 1799–2026Parsed source records: 1,674
Historical periods
Period
Years
Records
British Period
1799–1947
245
Pakistan Period
1948–1971
160
Bangladesh Period
1972–2026
1269
Total
1799–2026
1,674
Important… See the full description on the dataset page: https://huggingface.co/datasets/afzal-hosen-mandal/index-of-the-bangladesh-code.BanglaSleep-CoT
BanglaSleep-CoT
The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces.
Built for the Uncharted Data Challenge by Adaption Labs.
Expanded using Adaptive Data by Adaption.
Dataset at a Glance
Why This Dataset Exists
Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.dart_math_banglaThe dataset contains math problems in bangla. hkust-nlp/dart-math-uniform is translated using facebook/nllb-200-3.3B. To achive better performance english sentences are splitted and then fed into the translation model.
hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.JobCCC-Conversational-Job-Recommendation-Bangladesh
JobCCC: A Conversational Code-Mixed Corpus for Job Recommendation in Bangladesh
Dataset Creators
Authors: Md. Arman Hossain, Mubashir Jawad, Fariha Khandakar Moon, and Sonia Binte Siraj
Supervisor: Dr. Nafis Sadeq
Institution: Department of Computer Science & Engineering, East West University
Dataset Summary
JobCCC (Conversational Code-Mixed Corpus) is a multi-turn conversational benchmark and job recommendation dataset tailored for the… See the full description on the dataset page: https://huggingface.co/datasets/Armans33115/JobCCC-Conversational-Job-Recommendation-Bangladesh.banglabridge-instructions
Dataset Card — BanglaBridge Banglish Instruction Set
Summary
An original instruction-tuning dataset for code-mixed / romanized Bengali
("Banglish") — the register 100M+ people actually type online
(e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by
safe, deterministic transformation of our own templates. Nothing is scraped, so the
whole set is free to redistribute on Hugging Face and Kaggle.
This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.jugantor.com-scrape-bangla
Jugantor News Archive (Bangla)
Overview
This repository contains a comprehensive text dataset scraped from jugantor.com, one of the leading Bengali daily newspapers in Bangladesh. The primary goal of this archive is to preserve a massive collection of purely human-written journalism, editorials, and news reports, creating a distinct record of human-authored text separate from AI-generated content.
Purpose and Usage
This dataset is published publicly and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/jugantor.com-scrape-bangla.DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-DataDeepSeek R1 Bangla MMLU Distil Dataset
Original Dataset: hishab/bangla-mmlu
Train Samples: 17,796
Test Samples: 2,576
Total API Cost: 7K BDT
Contributors:
Myself
Numaer
How the Dataset was created
Step 1 - Base Dataset
I've used bangla-mmlu dataset released by hisab. Kudos to them for creating and open sourcing the dataset. Without their dataset this synthetic reasoning dataset won't exist in the first place.
Step 2 - Select Subset
Since I'm… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-Data.hsc-biology-bangla-dataset
🌿 HSC Biology Bangla Dataset (Plant Physiology)
The Ultimate Resource for Bengali STEM NLP
This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants.
✨ Key Highlights
Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.bangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.bangladesh-bar-council-exam-dataset
Bangladesh Bar Council Exam QA Dataset: 2022-2023 Bangla-English
The Bangladesh Bar Council Exam QA Dataset is a 400-question bilingual
Bangla-English legal question-answering benchmark compiled from the 2022 and
2023 Bangladesh Bar Council examination materials. It is intended for legal
NLP evaluation, multiple-choice question answering, retrieval experiments, and
Bangladesh law language-model research.
This repository contains the examination benchmark only. The larger… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-bar-council-exam-dataset.reveal-bangla
Reveal-Bangla:
Intro
Contains the Bangla translation of the subset from the reveal dataset.
Please refer to the following code snippet which has been used to select the subset:
SELECT *
FROM eval
Where ( answer_model = 'Flan-UL2-20B' or answer_model = 'GPT-3'
AND
answer_is_fully_attributable_and_correct = TRUE );
Only the following columns has been translated for the sake of the task:
question
full_answer
step
evidence
Usage
To load the dataset:
! pip… See the full description on the dataset page: https://huggingface.co/datasets/khondoker/reveal-bangla.bangla-nsfw-stories-scrape
Bangla Erotic Story Scraped Dataset
A comprehensive text dataset containing an archive of Bengali 18+ literature and stories scraped from various online sources.
📌 Dataset Summary
This repository serves as a text dataset aggregating Bengali adult fiction. The data has been collected to preserve these stories and can be used for natural language processing (NLP) tasks, linguistic analysis of colloquial Bengali, generative text modeling, or archival purposes.
⚠️… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-nsfw-stories-scrape.BanglaRQABanglaRQA is a human-annotated Bangla Question Answering (QA) dataset with diverse question-answer types.Bangladeshi_National_E-Services
Bangladeshi_National_E-Services Dataset
This Dataset contains all verified and authorized National e-services information of Bangladeshi Government
Description
I have collected these all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Bangladeshi_National_E-Services.BanglaSocialBench
BanglaSocialBench
Paper:BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction
Links
📄 Paper: https://aclanthology.org/2026.acl-srw.22/
💻 Code: https://github.com/sijantanvir/bangla-social-bench
Citation
@inproceedings{sijan-etal-2026-banglasocialbench,
title = {BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social… See the full description on the dataset page: https://huggingface.co/datasets/sijantanvir/BanglaSocialBench.Bangladeshi_Doctor_List
Bangladeshi_Doctor_List Dataset
This Dataset contains all verified and authorized Docto information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact
mahadise01@gmail.com
Linkdin:… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Bangladeshi_Doctor_List.bangla-bcs-qsBanglaBoroBangla_Education_Specialist_v1_11K
Bangla Education Specialist v1 — 11K
Dataset Description
A Bangla education-domain instruction-following dataset containing 11,000+ samples designed for fine-tuning Large Language Models (LLMs) on Bengali educational question-answering tasks.
Curated and processed by TechOptions.
Dataset Details
Property
Value
Language
Bengali (bn)
Domain
Education
Total Samples
~11,000
File Size
~6.48 MB
Format
JSONL → Parquet
License
Apache… See the full description on the dataset page: https://huggingface.co/datasets/techoptions/Bangla_Education_Specialist_v1_11K.
