datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BAD-Bengali-Aggressive-Text-Dataset
Novel Aggressive Text Dataset in Bengali
Tackling Cyber-Aggression: Identification and Fine-Grained Categorization of Aggressive Texts on Social Media using Weighted Ensemble of Transformers
Author: Omar Sharif and Mohammed Moshiul Hoque
Related Papers:
Paper1 in Neurocomputing Journal
Paper2 in CONSTRAINT@AAAI-2021
Paper3 in LTEDI@EACL-2021
Abstract
The pervasiveness of aggressive content in social media has become a serious concern for government… See the full description on the dataset page: https://huggingface.co/datasets/omar-sharif/BAD-Bengali-Aggressive-Text-Dataset.bengali-talkshow-audio
Bengali Talkshow Audio Dataset
A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs.
Dataset Description
This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.bengali-diarization-synthetic-v3shangkhachil-bengali-public-domain
Bengali Public-Domain Literature
101 complete works by 21 authors,
11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09.
Where these texts are read
https://shangkhachil.com — the reading site this corpus was built for. Free, no
account, 246 works by 28 authors. The complete text of
every work in this file can be read there.
This file is the text. The site is the part a JSONL cannot be:
Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.Bengali-PDbengali_sentimentben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.badlad-results
BaDLAD — public results
Score rows for the maintainer-run BaDLAD leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights
mask_map
COCO mask AP@[.5:.95] on private paper hidden test (higher better)
mask_map_by_domain
JSON map domain → mask_map
bbox_map
COCO bbox AP@[.5:.95] (optional)
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/BADLAD
evaluated_at
ISO date
requested_by
Who asked, or… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/badlad-results.BengaliMoralBench
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
Accepted at ACM FAccT 2026 · Montreal, QC, Canada · June 25–28, 2026
View in arXiv: https://arxiv.org/abs/2511.03180
View website: https://ciol-researchlab.github.io/works/BengaliMoralBench/
📋 Overview
BengaliMoralBench is the first large-scale, culturally grounded ethics benchmark for evaluating moral reasoning in Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/BengaliMoralBench.bengali_sentiment_v2@INPROCEEDINGS{10402353,
author={Hossain, Md Mehrab and Akhi, Iffat Ara Helal and Raisa, Syeda Rafiatus Sama and Onni, Taskin Sultana and Rifat, Ariful Islam and Islam, Ashraful},
booktitle={2023 IEEE 15th International Conference on Computational Intelligence and Communication Networks (CICN)},
title={Critical Analysis of BERT and LSTM Model for Bengali Sentiment Analysis Across Varied Datasets},
year={2023},
volume={},
number={},
pages={663-668},
keywords={Analytical… See the full description on the dataset page: https://huggingface.co/datasets/mHossain/bengali_sentiment_v2.Bengali-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Bengali STEM textbook data, containing 308 books and 12.88 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Bengali.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali-STEM-Textbook-Dataset.audio_tts_description_bengaliBengali-Non-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Bengali Non-STEM textbook data, containing 1,392 books and 74.34 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Bengali.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali-Non-STEM-Textbook-Dataset.bengaliai-competition-features-embeddings
BengaliAI Competition Embeddings and Features
bengali-diarization-synthetic-v4
Bengali Speaker Diarization Synthetic Dataset V4
Synthetic Bengali speaker diarization dataset with natural overlapping speech patterns using timeline-based random chunk placement.
Dataset Overview
Property
Value
Total Samples
600
Speaker Categories
1-30 speakers per sample
Samples per Category
20
Duration per Sample
~30 minutes
Total Duration
~300 hours
Sample Rate
16000 Hz
Format
WAV (audio) + RTTM (labels) + JSON (metadata)… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-diarization-synthetic-v4.BengaliFigbengali-tokenization-corpus
Bengali Tokenization Corpus (25k Sentences)
Dataset Description
A balanced 25,000-sentence Bengali corpus designed for tokenization benchmarking.
Dataset Summary
This dataset is used in the manuscript:
Quantifying the Tokenization Tax on Bengali: A Multi-Tokenizer Audit Across 25k Sentences
MD. Muktadirul Haque Maruf, Israt Jahan Munny, Moutithi Sarker MouManuscript in preparation
Domains
Academic
News
Literary
Colloquial
Dialectal… See the full description on the dataset page: https://huggingface.co/datasets/M-H-MARUF/bengali-tokenization-corpus.
