CoolFace
Datasetpublic

nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired

BLUGE-BCLAU: Bangla Linguistic Acceptability (Unpaired) BLUGE-BCLAU is a meticulously constructed Bangla Corpus of Linguistic Acceptability (Unpaired), one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it. Dataset Description This binary… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes36downloads
Dataset Card

BLUGE-BCLAU: Bangla Linguistic Acceptability (Unpaired)

BLUGE-BCLAU is a meticulously constructed Bangla Corpus of Linguistic Acceptability (Unpaired), one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it.

Dataset Description

This binary classification task complements **BLUGE-BCLAP** by evaluating grammatical acceptability judgment in a non-paired format. Unlike BCLAP, which contains both the correct and incorrect version of the same sentence, BCLAU consists of independently sampled grammatically acceptable and unacceptable Bangla sentences, drawn from the **BanglaGEC** corpus.

BCLAP and BCLAU are built from mutually exclusive samples — no sentence instance is shared between the two datasets — so each is designed to support a distinct assessment strategy. This unpaired format enables broader grammatical assessment by incorporating a diverse range of acceptable sentences that are not directly tied to specific, paired error corrections.

Dataset Structure

Balanced 80 / 10 / 10 split:

SplitSamples
Train102,799
Validation12,849
Test12,850
Total128,498

Fields:

  • —sentence — the Bangla sentence (ungrammatical or corrected)
  • —label — acceptability (`0`: unacceptable, `1`: acceptable)
  • —ErrorType — the grammatical error category (e.g. verb inflection, misplaced punctuation, missing subject)

Usage

python
from datasets import load_dataset

# Load all splits
ds = load_dataset("nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired")

# Access a specific split
train_ds = ds["train"]
val_ds   = ds["validation"]
test_ds  = ds["test"]

print(train_ds[0])

Load a single split directly:

python
from datasets import load_dataset

test_ds = load_dataset("nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired", split="test")

Or read the Parquet files directly with pandas:

python
import pandas as pd

train_df = pd.read_parquet(
    "hf://datasets/nahid-hub/BLUGE-bengali-linguistic-acceptability-unpaired/train.parquet"
)

Related Dataset

BCLAU is complemented by **BLUGE-BCLAP**, which offers the same acceptability judgment task in a paired format, linking each ungrammatical sentence directly to its corrected counterpart.

Source Data & Attribution

Built from the BanglaGEC corpus. Please refer to BanglaGEC's original source and license terms in addition to this dataset's license.

License

Released under CC BY 4.0.

Citation

If you use this dataset, please cite:

bibtex
@ARTICLE{BnLM-BLUGE-B-CORE,
  author={Hossain, Nahid and Faisal Kabir, Md.},
  journal={IEEE Access}, 
  title={Efficient Monolingual Pretraining in Low-Resource Settings Through Morphology-Aware Tokenization, Principled Corpus Denoising, and Benchmark-Driven Evaluation}, 
  year={2026},
  volume={14},
  number={},
  pages={91979-92003},
  keywords={Modeling;Multilingual;Training;Cleaning;Vocabulary;Labeling;Tokenization;Computational linguistics;Pipelines;Conferences;B-CORE;BLUGE;BnLM;corpus;evaluation benchmark;pretrained models},
  doi={10.1109/ACCESS.2026.3701520}
}