nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired
BLUGE-BCLAU: Bangla Linguistic Acceptability (Unpaired) BLUGE-BCLAU is a meticulously constructed Bangla Corpus of Linguistic Acceptability (Unpaired), one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it. Dataset Description This binary… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired.
BLUGE-BCLAU: Bangla Linguistic Acceptability (Unpaired)
BLUGE-BCLAU is a meticulously constructed Bangla Corpus of Linguistic Acceptability (Unpaired), one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it.
Dataset Description
This binary classification task complements **BLUGE-BCLAP** by evaluating grammatical acceptability judgment in a non-paired format. Unlike BCLAP, which contains both the correct and incorrect version of the same sentence, BCLAU consists of independently sampled grammatically acceptable and unacceptable Bangla sentences, drawn from the **BanglaGEC** corpus.
BCLAP and BCLAU are built from mutually exclusive samples — no sentence instance is shared between the two datasets — so each is designed to support a distinct assessment strategy. This unpaired format enables broader grammatical assessment by incorporating a diverse range of acceptable sentences that are not directly tied to specific, paired error corrections.
Dataset Structure
Balanced 80 / 10 / 10 split:
Fields:
sentence— the Bangla sentence (ungrammatical or corrected)label— acceptability (`0`: unacceptable, `1`: acceptable)ErrorType— the grammatical error category (e.g. verb inflection, misplaced punctuation, missing subject)
Usage
from datasets import load_dataset
# Load all splits
ds = load_dataset("nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired")
# Access a specific split
train_ds = ds["train"]
val_ds = ds["validation"]
test_ds = ds["test"]
print(train_ds[0])Load a single split directly:
from datasets import load_dataset
test_ds = load_dataset("nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired", split="test")Or read the Parquet files directly with pandas:
import pandas as pd
train_df = pd.read_parquet(
"hf://datasets/nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired/train.parquet"
)Related Dataset
BCLAU is complemented by **BLUGE-BCLAP**, which offers the same acceptability judgment task in a paired format, linking each ungrammatical sentence directly to its corrected counterpart.
Source Data & Attribution
Built from the BanglaGEC corpus. Please refer to BanglaGEC's original source and license terms in addition to this dataset's license.
License
Released under CC BY 4.0.
Citation
If you use this dataset, please cite:
@ARTICLE{BnLM-BLUGE-B-CORE,
author={Hossain, Nahid and Faisal Kabir, Md.},
journal={IEEE Access},
title={Efficient Monolingual Pretraining in Low-Resource Settings Through Morphology-Aware Tokenization, Principled Corpus Denoising, and Benchmark-Driven Evaluation},
year={2026},
volume={14},
number={},
pages={91979-92003},
keywords={Modeling;Multilingual;Training;Cleaning;Vocabulary;Labeling;Tokenization;Computational linguistics;Pipelines;Conferences;B-CORE;BLUGE;BnLM;corpus;evaluation benchmark;pretrained models},
doi={10.1109/ACCESS.2026.3701520}
}