datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BAD-Bengali-Aggressive-Text-Dataset
Novel Aggressive Text Dataset in Bengali
Tackling Cyber-Aggression: Identification and Fine-Grained Categorization of Aggressive Texts on Social Media using Weighted Ensemble of Transformers
Author: Omar Sharif and Mohammed Moshiul Hoque
Related Papers:
Paper1 in Neurocomputing Journal
Paper2 in CONSTRAINT@AAAI-2021
Paper3 in LTEDI@EACL-2021
Abstract
The pervasiveness of aggressive content in social media has become a serious concern for government… See the full description on the dataset page: https://huggingface.co/datasets/omar-sharif/BAD-Bengali-Aggressive-Text-Dataset.BHM-Bengali-Hateful-Memes
Dataset Description
BHM is a novel multimodal dataset for Bengali Hateful Memes detection. The dataset consists of 7,148 memes with Bengali as well as code-mixed captions,
tailored for two tasks: (i) detecting hateful memes and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).
Paper Information
Paper: https://aclanthology.org/2024.acl-long.454/
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Eftekhar/BHM-Bengali-Hateful-Memes.bengali-math-cotBengali-Fake-Review-DatasetThis is a binary dataset used for Bengali fake review detection in the paper "Bengali Fake Reviews: A Benchmark Dataset and Detection System" accepted in Neurocomputing, a journal published by Elsevier.
Annotated by 4 native Bangla speakers with more than 90% trustworthiness score.
Fleiss' Kappa Score: 0.83
Number of Taotal Data
Fake - 1339
Non-fake - 7710
Class wise statistics of BFRD dataset
Statistics
Fake
Non-fake
Total words
1,55,789
9,27,902
Total… See the full description on the dataset page: https://huggingface.co/datasets/shawon95/Bengali-Fake-Review-Dataset.bengali_sentimentwonbias-partial-dataset
Bengali Gender Bias Dataset: A Balanced Corpus for Analysis and Mitigation
Dataset Details
Overview
A manually annotated corpus of Bangla designed to benchmark and mitigate gender bias specifically targeting women. The dataset supports research in ethical NLP, hate speech detection, and computational social science.
Dataset Description
Basic Info
Purpose: Detect gender bias against women in Bengali text
Language: Bengali (Bangla)
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/gender-bias-bengali/wonbias-partial-dataset.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.BengaliMoralBench
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
Accepted at ACM FAccT 2026 · Montreal, QC, Canada · June 25–28, 2026
View in arXiv: https://arxiv.org/abs/2511.03180
View website: https://ciol-researchlab.github.io/works/BengaliMoralBench/
📋 Overview
BengaliMoralBench is the first large-scale, culturally grounded ethics benchmark for evaluating moral reasoning in Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/BengaliMoralBench.badlad-results
BaDLAD — public results
Score rows for the maintainer-run BaDLAD leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights
mask_map
COCO mask AP@[.5:.95] on private paper hidden test (higher better)
mask_map_by_domain
JSON map domain → mask_map
bbox_map
COCO bbox AP@[.5:.95] (optional)
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/BADLAD
evaluated_at
ISO date
requested_by
Who asked, or… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/badlad-results.bengali-names-vs-gender
Bengali Female VS Male Names Dataset
An NLP dataset that contains 2030 data samples of Bengali names and corresponding gender both for female and male. This is a very small and simple toy dataset that can be used by NLP starters to practice sequence classification problem and other NLP problems like gender recognition from names.
Background
In Bengali language, name of a person is dependent largely on their gender. Normally, name of a female ends with certain type of… See the full description on the dataset page: https://huggingface.co/datasets/faruk/bengali-names-vs-gender.bengali_sentiment_v2@INPROCEEDINGS{10402353,
author={Hossain, Md Mehrab and Akhi, Iffat Ara Helal and Raisa, Syeda Rafiatus Sama and Onni, Taskin Sultana and Rifat, Ariful Islam and Islam, Ashraful},
booktitle={2023 IEEE 15th International Conference on Computational Intelligence and Communication Networks (CICN)},
title={Critical Analysis of BERT and LSTM Model for Bengali Sentiment Analysis Across Varied Datasets},
year={2023},
volume={},
number={},
pages={663-668},
keywords={Analytical… See the full description on the dataset page: https://huggingface.co/datasets/mHossain/bengali_sentiment_v2.Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words
If you use Vulgar Lexicon dataset, please cite the following paper:
@Article{app132111875,
AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito},
TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla},
JOURNAL = {Applied Sciences},
VOLUME = {13},
YEAR = {2023},
NUMBER = {21},
ARTICLE-NUMBER = {11875},
URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/kit-nlp/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.bengali_sa
Sentiment Analysis Data for the Bengali Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Sazzed et al. (2020).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{sazzed-2020-cross,
title = "Cross-lingual sentiment classification in low-resource {B}engali language",
author = "Sazzed, Salim",
editor = "Xu, Wei and
Ritter, Alan… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/bengali_sa.bengali-hate-speech-datasetfrom https://www.kaggle.com/datasets/naurosromim/bengali-hate-speech-dataset?resource=download
just prefer to use through hf
Motamot_Bengali_Political_Sentiment_Analysis
Motamot: Bengali Political Sentiment Analysis Dataset
📖 Overview
Motamot is a Bengali political sentiment analysis dataset containing 7,058 labeled data points. Each entry is annotated with Positive or Negative sentiment, specifically tailored for analyzing political discourse in the Bengali language.
This dataset supports Natural Language Processing (NLP) research, with applications in sentiment classification, political opinion mining, and benchmarking pre-trained and… See the full description on the dataset page: https://huggingface.co/datasets/Mukaffi28/Motamot_Bengali_Political_Sentiment_Analysis.bengali-visual-genome-instruction-setopus100-Bengali-to-Englishbengali_deed_summarizewonbias-complete-dataset
WoNBias: A Bengali Dataset for Gender Bias Detection
Dataset Details
Overview
A manually annotated corpus of Bengali designed to identify gender-based biases, stereotypes, and harmful language against women. Supports research in ethical NLP, content moderation, and computational social science.
Dataset Description
Basic Info
Purpose: Detect gender bias and harmful language against women in Bengali text
Language: Bengali (Bangla)
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/gender-bias-bengali/wonbias-complete-dataset.Bengali_ParaPhrase_Corpus
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
a---
# For reference on dataset card metadata, see the spec: https://github.com/huggingface/hub-docs/blob/main/datasetcard.md?plain=1
# Doc / guide: https://huggingface.co/docs/hub/datasets-cards
Dataset Card for BanglaParaCorpus
BanglaParaCorpus is a large-scale Bengali paraphrase generation… See the full description on the dataset page: https://huggingface.co/datasets/mHossain/Bengali_ParaPhrase_Corpus.bengali_chat_conv
Bengali Chat Conversation Dataset
Dataset Overview
The Bengali Chat Conversation dataset contains a collection of conversational exchanges in Bengali. Each entry consists of a question and its corresponding answer, covering a wide range of topics including technology, health, environment, education, and more. This dataset can be used for various natural language processing (NLP) tasks such as language modeling, question-answering systems, and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/spedrox-sac/bengali_chat_conv.bengali-words
Bengali Words
This dataset contains a Bengali/Bangla word list.
Dataset Format
The dataset is provided as a CSV file. Words are sorted in alphabetical order.
word
অগ্রদানী
অধ্যাত্মশক্তি
অননুসরণ
অবলীলাক্রমে
অরুণরঞ্জিত
...
Columns
Column
Description
word
Bengali word
Installation
pip install -U datasets
Usage
from datasets import load_dataset
dataset = load_dataset("kawshikbuet17/bengali-words")… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-words.bengali-poemsVulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words
If you use Vulgar Lexicon dataset, please cite the following paper:
@Article{app132111875,
AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito},
TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla},
JOURNAL = {Applied Sciences},
VOLUME = {13},
YEAR = {2023},
NUMBER = {21},
ARTICLE-NUMBER = {11875},
URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/TanjimKIT/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.bengali-hindi-number-blindspot
Blind Spots of Frontier Models: Bengali & Hindi Number Word-to-Digit Conversion
Summary
This dataset documents a critical blind spot in small open-source language models: failure to correctly convert Bengali and Hindi number words into their digit equivalents. Bengali and Hindi share the South Asian number system (hazar/হাজার, lakh/লাখ, crore/কোটি), and all three tested models consistently fail at this fundamental conversion step.
IMPORTANT: Arithmetic calculation errors… See the full description on the dataset page: https://huggingface.co/datasets/wong132/bengali-hindi-number-blindspot.Bengali-WordlistBengali_sum_v1BengaliQADatasetBengaliNewsRoleplay-Bengali
RolePlay-Bengali
Roleplay-Bengali Dataset is a dataset for roleplaying in the Bengali language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Bengali.
