CoolFace
Datasetpublicgated

nahidstaq/bangla-llm-data

Bangla NLP Text Corpus — 800K+ Bangla Text Samples for LLM Training and NLP Research The largest open, multi-domain Bangla text dataset, combining 801,645 samples from 15 different sources — newspapers, social media, education, reviews, QA, medical, poetry, and more. Ready for Bangla LLM pretraining, fine-tuning, and downstream NLP tasks. Why This Dataset? Bangla (Bengali) is the 7th most spoken language in the world with 230M+ speakers, but high-quality Bangla… See the full description on the dataset page: https://huggingface.co/datasets/nahidstaq/bangla-llm-data.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
2likes24downloads

nahidstaq/bangla-llm-data · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.