nahid-hub/BLUGE-bengali-news-classification
BLUGE-NCC: Bangla News Classification BLUGE-NCC is a meticulously curated and balanced Bangla News Category Classification dataset, one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it. Dataset Description This task classifies Bangla news articles… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-news-classification.
BLUGE-NCC: Bangla News Classification
BLUGE-NCC is a meticulously curated and balanced Bangla News Category Classification dataset, one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it.
Dataset Description
This task classifies Bangla news articles into seven categories:
Dataset Structure
Balanced 80 / 10 / 10 split:
Fields:
text— the Bangla news article contentlabel— news category (0–6, see table above)
Usage
from datasets import load_dataset
# Load all splits
ds = load_dataset("nahid-hub/BLUGE-bengali-news-classification")
# Access a specific split
train_ds = ds["train"]
val_ds = ds["validation"]
test_ds = ds["test"]
print(train_ds[0])Load a single split directly:
from datasets import load_dataset
test_ds = load_dataset("nahid-hub/BLUGE-bengali-news-classification", split="test")Or read the Parquet files directly with pandas:
import pandas as pd
train_df = pd.read_parquet(
"hf://datasets/nahid-hub/BLUGE-bengali-news-classification/train.parquet"
)License
Released under CC BY 4.0 for the dataset compilation, labels, and splits. You are free to share and adapt this dataset for research and most other purposes, provided you give appropriate credit — see the note above regarding underlying article copyright.
Citation
If you use this dataset, please cite:
@ARTICLE{BnLM-BLUGE-B-CORE,
author={Hossain, Nahid and Faisal Kabir, Md.},
journal={IEEE Access},
title={Efficient Monolingual Pretraining in Low-Resource Settings Through Morphology-Aware Tokenization, Principled Corpus Denoising, and Benchmark-Driven Evaluation},
year={2026},
volume={14},
number={},
pages={91979-92003},
keywords={Modeling;Multilingual;Training;Cleaning;Vocabulary;Labeling;Tokenization;Computational linguistics;Pipelines;Conferences;B-CORE;BLUGE;BnLM;corpus;evaluation benchmark;pretrained models},
doi={10.1109/ACCESS.2026.3701520}
}