CoolFace
Datasetpublic

rejauldu/bengali-wikipedia

📚 Bengali Wikipedia Language Modeling Dataset (For GPT-2 Training) 📝 Dataset Summary This dataset contains a large Bengali text corpus collected from Bengali Wikipedia.It is cleaned, sentence-segmented, and formatted for next-token prediction language modeling tasks such as GPT-2 training. It includes train and validation splits, suitable for transformer-based Bengali language models. 📊 Dataset Details Property Value Language Bengali… See the full description on the dataset page: https://huggingface.co/datasets/rejauldu/bengali-wikipedia.

sourceHugging Facecc-by-3.0updated 11mo agoView on Hugging Face
0likes70downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
rejauldu/bengali-wikipedia · CoolFace