CoolFace
Datasetpublic

nahid-hub/BLUGE-bengali-news-classification

BLUGE-NCC: Bangla News Classification BLUGE-NCC is a meticulously curated and balanced Bangla News Category Classification dataset, one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it. Dataset Description This task classifies Bangla news articles… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-news-classification.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes90downloads
Dataset Card

BLUGE-NCC: Bangla News Classification

BLUGE-NCC is a meticulously curated and balanced Bangla News Category Classification dataset, one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it.

Dataset Description

This task classifies Bangla news articles into seven categories:

LabelCategory
0Business
1Technology
2Crime
3Entertainment
4International Affairs
5Sports
6Lifestyle

Dataset Structure

Balanced 80 / 10 / 10 split:

SplitSamples
Train49,840
Validation6,230
Test6,230
Total62,300

Fields:

  • —text — the Bangla news article content
  • —label — news category (0–6, see table above)

Usage

python
from datasets import load_dataset

# Load all splits
ds = load_dataset("nahid-hub/BLUGE-bengali-news-classification")

# Access a specific split
train_ds = ds["train"]
val_ds   = ds["validation"]
test_ds  = ds["test"]

print(train_ds[0])

Load a single split directly:

python
from datasets import load_dataset

test_ds = load_dataset("nahid-hub/BLUGE-bengali-news-classification", split="test")

Or read the Parquet files directly with pandas:

python
import pandas as pd

train_df = pd.read_parquet(
    "hf://datasets/nahid-hub/BLUGE-bengali-news-classification/train.parquet"
)

License

Released under CC BY 4.0 for the dataset compilation, labels, and splits. You are free to share and adapt this dataset for research and most other purposes, provided you give appropriate credit — see the note above regarding underlying article copyright.

Citation

If you use this dataset, please cite:

bibtex
@ARTICLE{BnLM-BLUGE-B-CORE,
  author={Hossain, Nahid and Faisal Kabir, Md.},
  journal={IEEE Access}, 
  title={Efficient Monolingual Pretraining in Low-Resource Settings Through Morphology-Aware Tokenization, Principled Corpus Denoising, and Benchmark-Driven Evaluation}, 
  year={2026},
  volume={14},
  number={},
  pages={91979-92003},
  keywords={Modeling;Multilingual;Training;Cleaning;Vocabulary;Labeling;Tokenization;Computational linguistics;Pipelines;Conferences;B-CORE;BLUGE;BnLM;corpus;evaluation benchmark;pretrained models},
  doi={10.1109/ACCESS.2026.3701520}
}