CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01md-nishat-008 /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.texttext-generation10K<n<100K2 likes148 downloads1y agoHugging Face02md-nishat-008 /Bangla-Instruct Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.texttext-generation100K<n<1M8 likes119 downloads1y agoHugging Face03Mahadih534 /Institutional-Information-of-Bangladesh Institutional-Information-of-Bangladesh Dataset This Dataset contains all verified and authorized Institutional information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.tabularquestion-answering10K<n<100K2 likes114 downloads2y agoHugging Face04mehedihasanbijoy /BanglaSEC BanglaSEC A 1.18M-pair parallel corpus for Bangla spelling error correction, with character-level error masks across 14 error types. BanglaSEC is the corpus introduced in A transformer based spelling error correction framework for Bangla and resource scarce Indic languages (Bijoy, Hossain, Islam & Shatabda, Computer Speech & Language 89:101703, 2025). Each row pairs a correct Bangla word with an erroneous form, labelled by error type and annotated with a binary mask marking… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaSEC.tabulartext-generation1M<n<10M0 likes106 downloads13d agoHugging Face05tasfuuu19 /BanglaSleep-CoT BanglaSleep-CoT The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces. Built for the Uncharted Data Challenge by Adaption Labs. Expanded using Adaptive Data by Adaption. Dataset at a Glance Why This Dataset Exists Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.tabulartext-generation1K<n<10K0 likes71 downloads5mo agoHugging Face06nahid-hub /BanglaGEC BanglaGEC: A Large-Scale Parallel Corpus for Bangla Grammatical Error Correction BanglaGEC is a large-scale parallel corpus of 7,074,425 (~7.1M) sentence pairs for Bangla (Bengali) Grammatical Error Correction (GEC). Each pair maps a grammatically erroneous Bangla sentence to its grammatically correct counterpart, along with the error type, making it directly usable for training and evaluating sequence-to-sequence models, transformers, and large language models on the Bangla… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BanglaGEC.texttext-generation1M<n<10M0 likes66 downloads2mo agoHugging Face07mehedihasanbijoy /BanglaPRCorpus BanglaPRCorpus A 1.48M-pair corpus for Bangla punctuation restoration — unpunctuated source sentences paired with their fully punctuated targets, labelled by how many punctuation marks were removed. BanglaPRCorpus is the corpus introduced in Advancing Bangla Punctuation Restoration by a Monolingual Transformer-Based Method and a Large-Scale Corpus (Bijoy et al., EMNLP 2023 Workshop on Bangla Language Processing), alongside the Jatikarok model. Each row is a (source, target)… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaPRCorpus.texttext-generation1M<n<10M0 likes58 downloads13d agoHugging Face08faisal4590aziz /bangla-health-related-paraphrased-dataset Dataset Card for "BanglaHealthParaphrase" BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.tabulartext-generation100K<n<1M2 likes30 downloads1y agoHugging Face09md-nishat-008 /Bangla-Code-Instruct 🐯 Bangla-Code-Instruct: A Comprehensive Bangla Code Instruction Dataset Accepted at LREC 2026 Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri George Mason University, Fairfax, VA, USA The first large-scale Bangla code instruction dataset (300K examples) for training Code LLMs in Bangla. ⚠️ Note: The dataset will be released after the LREC 2026 conference. Stay tuned! Overview Bangla-Code-Instruct is a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Code-Instruct.texttext-generation100K<n<1M0 likes26 downloads6mo agoHugging Face10midnightGlow /BanglaSumWe created a dataset by web scraping different online newspapers like ‘The Daily Star’, ‘ProthomAlo’, and ‘BBC News Bangla’ using the Beautiful Soup library of Python. The dataset's features are ‘title’, ‘text’ & ‘summary’. The dataset was preprocessed using the Python library ‘pandas’ and all duplicates and null values were eliminated and the total number of rows remaining is 9311. Since there is a lack of datasets in the Bengali language, this dataset will be useful for tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/midnightGlow/BanglaSum.textsummarization1K<n<10K0 likes20 downloads2y agoHugging Face11Mahadih534 /Bangladeshi_Doctor_List Bangladeshi_Doctor_List Dataset This Dataset contains all verified and authorized Docto information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact mahadise01@gmail.com Linkdin:… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Bangladeshi_Doctor_List.tabularquestion-answeringn<1K1 likes18 downloads2y agoHugging Face12BadarHossain /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/BadarHossain/Bangla-TextBook.texttext-generation10K<n<100K0 likes15 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.