datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangladesh-stock-market-dataset
Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks
Author: Kawser Sikder
Overview
A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market.
Metric
Value
Total Stocks
441
Total Sectors
23
Total Trading Records
1,507,388
Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.Bangla-TextBook
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
---
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.Bangla-Instruct
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.Institutional-Information-of-Bangladesh
Institutional-Information-of-Bangladesh Dataset
This Dataset contains all verified and authorized Institutional information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.bangla-emergency-posts
Bangla Emergency Posts
5,836 Bangla social media posts, hand-labelled into nine emergency categories.
Built for Bangla Emergency Post Classification on Social Media using Transformer
Based BERT Models (EICT 2023).
Emergency text classification in Bangla is scarce despite the language having
hundreds of millions of speakers. This dataset exists so that Bangla-speaking
people can report emergencies in their own language and have those reports
routed automatically.… See the full description on the dataset page: https://huggingface.co/datasets/NightRaven/bangla-emergency-posts.BanglaMedQA
Dataset Card for BanglaMedQA and BanglaMMedBench
This dataset introduces BanglaMedQA and BanglaMMedBench, the first large-scale Bangla biomedical Multiple Choice Question (MCQ) datasets designed to evaluate reasoning and retrieval-based methods such as Retrieval-Augmented Generation (RAG) for Bangla Question Answering.
Dataset Details
Dataset Description
BanglaMedQA consists of 1,000 MCQs collected from authentic Bangladeshi medical admission exams (MBBS, BDS… See the full description on the dataset page: https://huggingface.co/datasets/ajwad-abrar/BanglaMedQA.BanglaSEC
BanglaSEC
A 1.18M-pair parallel corpus for Bangla spelling error correction, with character-level error masks across 14 error types.
BanglaSEC is the corpus introduced in A transformer based spelling error correction framework for Bangla and resource scarce Indic languages (Bijoy, Hossain, Islam & Shatabda, Computer Speech & Language 89:101703, 2025). Each row pairs a correct Bangla word with an erroneous form, labelled by error type and annotated with a binary mask marking… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaSEC.Bangladeshi-Restaurant-Dataall-bangladeshi-hospitalsBanglaAirlineIntent
Bangla / English / Banglish Airline Support Intent Classification
A 19-intent classification dataset for an airline customer-support chatbot,
modelled on US-Bangla Airlines' passenger mix and covering the four ways those
passengers actually write:
script
example
rows
bn Bengali script
আমার ফ্লাইট কি সময়মতো ছাড়বে
3,174
en English
has BG147 landed yet
3,079
bl Banglish (romanized Bangla)
amar flight ta ki time mto charbe
3,062
mx code-mixed mid-sentence
BG435 ki… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaAirlineIntent.BanglaTLit
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla
Dataset Overview
BanglaTLit-PT: A pre-training corpus with 245727 transliterated or romanized Bangla samples for further pre-training language models.
BanglaTLit: Subset of the BanglaTLit-PT dataset containing 42705 romanized Bangla and its corresponding Bangla back-transliteration pairs.
Data Description
Column Title
Description
id
A unique identifier for each data point… See the full description on the dataset page: https://huggingface.co/datasets/aplycaebous/BanglaTLit.BanglaSleep-CoT
BanglaSleep-CoT
The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces.
Built for the Uncharted Data Challenge by Adaption Labs.
Expanded using Adaptive Data by Adaption.
Dataset at a Glance
Why This Dataset Exists
Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.BanglaGEC
BanglaGEC: A Large-Scale Parallel Corpus for Bangla Grammatical Error Correction
BanglaGEC is a large-scale parallel corpus of 7,074,425 (~7.1M) sentence pairs for Bangla (Bengali) Grammatical Error Correction (GEC). Each pair maps a grammatically erroneous Bangla sentence to its grammatically correct counterpart, along with the error type, making it directly usable for training and evaluating sequence-to-sequence models, transformers, and large language models on the Bangla… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BanglaGEC.all-Bangladeshi-medicinesBanglaPRCorpus
BanglaPRCorpus
A 1.48M-pair corpus for Bangla punctuation restoration — unpunctuated source sentences paired with their fully punctuated targets, labelled by how many punctuation marks were removed.
BanglaPRCorpus is the corpus introduced in Advancing Bangla Punctuation Restoration by a Monolingual Transformer-Based Method and a Large-Scale Corpus (Bijoy et al., EMNLP 2023 Workshop on Bangla Language Processing), alongside the Jatikarok model.
Each row is a (source, target)… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaPRCorpus.dart_math_banglaThe dataset contains math problems in bangla. hkust-nlp/dart-math-uniform is translated using facebook/nllb-200-3.3B. To achive better performance english sentences are splitted and then fed into the translation model.
bangla-phishing-detection-2026
Bangla Phishing Detection Dataset (SMS, Email, URLs) 2026
Synthetic dataset (~2000 rows) of phishing and legitimate messages in Bangla (Bengali) + some English, simulating common Bangladesh scams (bKash, Nagad, Daraz, Eid offers, job fraud, account lock alerts, etc.).
Research Motivation
Phishing/smishing attacks are rising in Bangladesh and South Asia, often in Bangla using local services. Most phishing datasets are English-only and miss these patterns.This dataset fills… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/bangla-phishing-detection-2026.Bangladesh_Locations
Bangladesh Postcodes Dataset (Bilingual & Structured)
A comprehensive, cleaned, and bilingual (English & Bangla) database of postal codes in Bangladesh. This dataset covers the full administrative hierarchy: Division > District > Thana (Upazila) > Post Office.
📂 Files Included
Filename
Format
Description
bangladesh_postcodes_final.csv
CSV
The master dataset with all columns. Best for data analysis or database imports.
bangladesh_postcodes_flat.json
JSON
A… See the full description on the dataset page: https://huggingface.co/datasets/ahnafch01/Bangladesh_Locations.Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP
Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP
Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.Bangladesh-Voter-Synthetic-Dataset
🗳️ Bangladesh Voter Dataset
📜 Dataset Description
The Bangladesh Voter Dataset is a synthetic dataset containing voter information for the purpose of demonstrating data generation and processing techniques. Each voter record includes both Bengali and English names, gender, NID, address, and profile information.
📊 Dataset Structure
The dataset is structured as follows:
profile: A URL to the voter's profile image.
nid: A unique National Identification Number.… See the full description on the dataset page: https://huggingface.co/datasets/jonybepary/Bangladesh-Voter-Synthetic-Dataset.bangla-health-related-paraphrased-dataset
Dataset Card for "BanglaHealthParaphrase"
BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.doctor_qa_banglaBangla-Math
Overview
The Bangla-Math Dataset is a valuable resource that addresses the critical need for Bangla-language mathematical problem-solving datasets. Currently, there are no publicly available datasets for math problems in Bangla, making this dataset a unique and valuable contribution to the fields of natural language processing (NLP). This dataset bridges the gap by enabling AI models to understand and work effectively with Bengali math content, thereby supporting advancements in… See the full description on the dataset page: https://huggingface.co/datasets/kawchar85/Bangla-Math.BanglaBookbangla-law-qnaBangla_Person_Name_ExtractorVulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words
If you use Vulgar Lexicon dataset, please cite the following paper:
@Article{app132111875,
AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito},
TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla},
JOURNAL = {Applied Sciences},
VOLUME = {13},
YEAR = {2023},
NUMBER = {21},
ARTICLE-NUMBER = {11875},
URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/kit-nlp/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.preprocessed-BanglaNMTreveal-bangla
Reveal-Bangla:
Intro
Contains the Bangla translation of the subset from the reveal dataset.
Please refer to the following code snippet which has been used to select the subset:
SELECT *
FROM eval
Where ( answer_model = 'Flan-UL2-20B' or answer_model = 'GPT-3'
AND
answer_is_fully_attributable_and_correct = TRUE );
Only the following columns has been translated for the sake of the task:
question
full_answer
step
evidence
Usage
To load the dataset:
! pip… See the full description on the dataset page: https://huggingface.co/datasets/khondoker/reveal-bangla.Bangla-Code-Instruct
🐯 Bangla-Code-Instruct: A Comprehensive Bangla Code Instruction Dataset
Accepted at LREC 2026
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri
George Mason University, Fairfax, VA, USA
The first large-scale Bangla code instruction dataset (300K examples) for training Code LLMs in Bangla.
⚠️ Note: The dataset will be released after the LREC 2026 conference. Stay tuned!
Overview
Bangla-Code-Instruct is a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Code-Instruct.
