datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BAD-Bengali-Aggressive-Text-Dataset
Novel Aggressive Text Dataset in Bengali
Tackling Cyber-Aggression: Identification and Fine-Grained Categorization of Aggressive Texts on Social Media using Weighted Ensemble of Transformers
Author: Omar Sharif and Mohammed Moshiul Hoque
Related Papers:
Paper1 in Neurocomputing Journal
Paper2 in CONSTRAINT@AAAI-2021
Paper3 in LTEDI@EACL-2021
Abstract
The pervasiveness of aggressive content in social media has become a serious concern for government… See the full description on the dataset page: https://huggingface.co/datasets/omar-sharif/BAD-Bengali-Aggressive-Text-Dataset.bengali_sentimentben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.BengaliMoralBench
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
Accepted at ACM FAccT 2026 · Montreal, QC, Canada · June 25–28, 2026
View in arXiv: https://arxiv.org/abs/2511.03180
View website: https://ciol-researchlab.github.io/works/BengaliMoralBench/
📋 Overview
BengaliMoralBench is the first large-scale, culturally grounded ethics benchmark for evaluating moral reasoning in Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/BengaliMoralBench.badlad-results
BaDLAD — public results
Score rows for the maintainer-run BaDLAD leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights
mask_map
COCO mask AP@[.5:.95] on private paper hidden test (higher better)
mask_map_by_domain
JSON map domain → mask_map
bbox_map
COCO bbox AP@[.5:.95] (optional)
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/BADLAD
evaluated_at
ISO date
requested_by
Who asked, or… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/badlad-results.bengali_sentiment_v2@INPROCEEDINGS{10402353,
author={Hossain, Md Mehrab and Akhi, Iffat Ara Helal and Raisa, Syeda Rafiatus Sama and Onni, Taskin Sultana and Rifat, Ariful Islam and Islam, Ashraful},
booktitle={2023 IEEE 15th International Conference on Computational Intelligence and Communication Networks (CICN)},
title={Critical Analysis of BERT and LSTM Model for Bengali Sentiment Analysis Across Varied Datasets},
year={2023},
volume={},
number={},
pages={663-668},
keywords={Analytical… See the full description on the dataset page: https://huggingface.co/datasets/mHossain/bengali_sentiment_v2.bengali-tokenization-corpus
Bengali Tokenization Corpus (25k Sentences)
Dataset Description
A balanced 25,000-sentence Bengali corpus designed for tokenization benchmarking.
Dataset Summary
This dataset is used in the manuscript:
Quantifying the Tokenization Tax on Bengali: A Multi-Tokenizer Audit Across 25k Sentences
MD. Muktadirul Haque Maruf, Israt Jahan Munny, Moutithi Sarker MouManuscript in preparation
Domains
Academic
News
Literary
Colloquial
Dialectal… See the full description on the dataset page: https://huggingface.co/datasets/M-H-MARUF/bengali-tokenization-corpus.
