datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangla-nlp-catalog
Bangla NLP Catalog
A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link.
This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them.
Why this exists
Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP
Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP
Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words
If you use Vulgar Lexicon dataset, please cite the following paper:
@Article{app132111875,
AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito},
TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla},
JOURNAL = {Applied Sciences},
VOLUME = {13},
YEAR = {2023},
NUMBER = {21},
ARTICLE-NUMBER = {11875},
URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/kit-nlp/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.BanglaNLP
BanglaNLP: Bengali-English Parallel Dataset Tools
BanglaNLP is a comprehensive toolkit for creating high-quality Bengali-English parallel datasets from news sources, designed to improve machine translation and other cross-lingual NLP tasks for the Bengali language. Our work addresses the critical shortage of high-quality parallel data for Bengali, the 7th most spoken language in the world with over 230 million speakers.
🏆 Impact & Recognition
120K+ Sentence Pairs:… See the full description on the dataset page: https://huggingface.co/datasets/likhonsheikh/BanglaNLP.
