datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
titulm-bangla-corpus
TituLM Bangla Corpus
This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability.
This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/hishab/titulm-bangla-corpus.BanglaEng-SynCorpus
BanglaEng-SynCorpus
Dataset Summary
BanglaEng-SynCorpus is a large-scale synthetic Bangla–English parallel corpus designed to support research in Neural Machine Translation (NMT) and other Bangla–English bilingual NLP tasks.The corpus is generated using linguistically validated sentence templates combined with topic-wise curated vocabularies, covering all 12 English/Bangla tense structures.
Due to extreme scale (trillions of possible sentence pairs), the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Eamin-sust/BanglaEng-SynCorpus.bangladesh-stock-market-dataset
Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks
Author: Kawser Sikder
Overview
A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market.
Metric
Value
Total Stocks
441
Total Sectors
23
Total Trading Records
1,507,388
Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.titulm-bangla-corpus
TituLM Bangla Corpus
This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability.
This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we… See the full description on the dataset page: https://huggingface.co/datasets/shofikul-1234/titulm-bangla-corpus.titulm-bangla-mmlu
Titulm Bangla MMLU
Read the paper for details: https://arxiv.org/abs/2502.11187
Citation
@misc{nahin2025titullmsfamilybanglallms,
title={TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking},
author={Shahriar Kabir Nahin and Rabindra Nath Nandi and Sagor Sarker and Quazi Sarwar Muhtaseem and Md Kowsher and Apu Chandraw Shill and Md Ibrahim and Mehadi Hasan Menon and Tareq Al Muntasir and Firoj Alam},
year={2025},
eprint={2502.11187}… See the full description on the dataset page: https://huggingface.co/datasets/hishab/titulm-bangla-mmlu.BanglaQwen-Train-Corpusbangla-corpus
BanglaBox — Bangladeshi Bangla TTS corpus
Anonymous artifact for double-blind review. A Bangladeshi Bangla speech corpus for text-to-speech and
zero-shot voice cloning, built with the coverage-driven script pipeline described in the paper
(7 domains — news, customer care, teaching, healthcare, e-commerce, finance, IT — with scripts selected under a
tiered Jensen–Shannon-divergence objective over phones, diphones, triphones and conjunct clusters
(juktakkhor) and filtered by… See the full description on the dataset page: https://huggingface.co/datasets/Banglabox/bangla-corpus.bangla-noise-robustness-databanglaTabQA
Dataset Card for "banglaTabQA"
Usage
import pandas as pd
from datasets import load_dataset
banglatableQA = load_dataset("vaishali/banglaTabQA")
for sample in banglatableQA['train']:
question = sample['question']
input_table = pd.read_json(sample['table'], orient='split')
answer = pd.read_json(sample['answer'], orient='split')
BibTeX entry and citation info
@inproceedings{pal-etal-2024-table,
title = "Table Question Answering for Low-resourced… See the full description on the dataset page: https://huggingface.co/datasets/vaishali/banglaTabQA.bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,816-hour Bengali speech corpus with
624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus
(567,323 recordings) and a separately collected 745.1-hour evaluation set
(57,628 recordings). It combines scripted single-speaker read speech with
natural multi-speaker conversations for Bengali automatic speech recognition
(ASR).
The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.ipfs_bangladesh_laws_ir
Bangladesh legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_bangladesh_laws (revision 16782096c126f7342b3cfeaa312c437c9fa2de73) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Bangladesh prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_bangladesh_laws_ir.bangla-crime-investigation-patterns-v2
Bangla Crime Investigation Patterns V2
This dataset is an anonymized, structured extraction from 1,078 public Bangla crime-investigation video transcripts. It was built to support machine-learning and LLM research on crime-pattern information extraction, case summarization, event sequencing, entity/relationship extraction, and Bengali investigative-report analysis.
The public release does not include raw full transcripts. It contains short redacted evidence snippets linked to… See the full description on the dataset page: https://huggingface.co/datasets/tanziro/bangla-crime-investigation-patterns-v2.BanglaRQABanglaRQA is a human-annotated Bangla Question Answering (QA) dataset with diverse question-answer types.bangladeshi-jobs
Bangladeshi Tech Jobs — Open Dataset
Weekly-refreshed, structured dataset of open software & IT job postings from Bangladeshi tech companies, built by an automated crawl → LLM-extraction → data-warehouse pipeline. Published as JSON + Parquet + a DuckDB star schema, free for any use with attribution (CC-BY-4.0).
Snapshot (2026-09-20)
🤖 Auto-generated on every build — these numbers are never edited by hand.
Metric
Value
Registered companies
234… See the full description on the dataset page: https://huggingface.co/datasets/swadhinbiswas/bangladeshi-jobs.bangla-newsbangla-instruction-dataset
🧠 Bangla Instruction Dataset
This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models.
📚 Dataset Splits
The dataset is organized into the following splits:
Split Name
Source Dataset
Description
OdiaGenAI
OdiaGenAI/all_combined_bengali_252k
A large-scale collection of diverse Bangla instructions and responses.
chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.BanglaVerse
Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects
Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision–language… See the full description on the dataset page: https://huggingface.co/datasets/FaiyazAbdullah114708/BanglaVerse.BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.bangla-voice-03042bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.translated_gsm8k_to_bangla_trainBanglaContextualBias
Dataset Card for Bangla Contextual Bias
The Bangla Contextual Bias dataset corresponds to the data described in the paper "An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla" accepted in ACL 2024 (Findings).
Dataset Description
The dataset has different parts for different bias detection experiments conducted for Bengali.
WEAT & SEAT
For the WEAT experiment, the dataset is translated from its English counterpart and… See the full description on the dataset page: https://huggingface.co/datasets/csebuetnlp/BanglaContextualBias.bangla-ocr-double-benchmark
Bangla OCR Double Benchmark
Two equally weighted, deterministic full-page Bangla handwriting robustness splits:
bongabdo: 6,669 readability-preserving renderings balanced over all 111 Bongabdo pages.
bn_htrd: 6,669 renderings balanced over all 75 actual files in the writer-separated
BN-HTRd test split.
These are explicitly compositional/augmentation robustness rows, not 13,338 independent
writers or source documents. Every row exposes its source page ID, source SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/bangla-ocr-double-benchmark.Bangla_MLM_Texts_DatasetBangla-TextBook
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
---
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.BanglaNMTThis is the largest Machine Translation (MT) dataset for Bengali-English, introduced in the paper
`Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation`.BanglaSafe
BanglaSafe dataset card
Overview
BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written
natively rather than translated from English. Every category is anchored to a Bangladesh statute or
a documented case, and every harm instance is written five ways so that only the language and the
register change.
That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.bangla-corpus
TituLM Bangla Corpus
This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability.
This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/munzurul/bangla-corpus.bangla_newspaper_dataset
Bangla Newspaper Dataset
400k+ bangla news samples, 25+ categories
Source
Data collected from https://www.prothomalo.com/archive [Copyright owned by the actual source]
Github
Github repository (Bi-LSTM Baseline): https://github.com/zabir-nabil/bangla-news-rnn
Kaggle Version
Kaggle Dataset: https://www.kaggle.com/datasets/furcifer/bangla-newspaper-dataset
Inspiration
The dataset can be used for Bangla text classification and generation… See the full description on the dataset page: https://huggingface.co/datasets/zabir-nabil/bangla_newspaper_dataset.bangla_tts_iitm
