CoolFace
Datasetpublic

nahid-hub/B-CORE-bengali-corpus

B-CORE: Bangla Pretraining Corpus B-CORE (Bengali Context-aware Optimized and Refined Entities) is a large-scale, rigorously curated Bangla monolingual corpus for language model pretraining, comprising 16.5 million documents (4.32 billion tokens, 52GB (20.8 GB Compressed)). It is among the largest and most carefully curated Bangla pretraining corpora available, constructed through a reproducible multi-stage pipeline. B-CORE was used to pretrain the BnLM-F and BnLM-C Bengali… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/B-CORE-bengali-corpus.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes107downloads

nahid-hub/B-CORE-bengali-corpus · main · files are served by the source, never re-hosted here