CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kawsersikder /bangladesh-stock-market-dataset Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks Author: Kawser Sikder Overview A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market. Metric Value Total Stocks 441 Total Sectors 23 Total Trading Records 1,507,388 Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.tabulartime-series-forecasting1M<n<10M1 likes1.1k downloads1mo agoHugging Face02md-nishat-008 /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.texttext-generation10K<n<100K2 likes148 downloads1y agoHugging Face03md-nishat-008 /Bangla-Instruct Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.texttext-generation100K<n<1M8 likes119 downloads1y agoHugging Face04Mahadih534 /Institutional-Information-of-Bangladesh Institutional-Information-of-Bangladesh Dataset This Dataset contains all verified and authorized Institutional information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.tabularquestion-answering10K<n<100K2 likes114 downloads2y agoHugging Face05NightRaven /bangla-emergency-posts Bangla Emergency Posts 5,836 Bangla social media posts, hand-labelled into nine emergency categories. Built for Bangla Emergency Post Classification on Social Media using Transformer Based BERT Models (EICT 2023). Emergency text classification in Bangla is scarce despite the language having hundreds of millions of speakers. This dataset exists so that Bangla-speaking people can report emergencies in their own language and have those reports routed automatically.… See the full description on the dataset page: https://huggingface.co/datasets/NightRaven/bangla-emergency-posts.texttext-classification10K<n<100K0 likes111 downloads21d agoHugging Face06ajwad-abrar /BanglaMedQA Dataset Card for BanglaMedQA and BanglaMMedBench This dataset introduces BanglaMedQA and BanglaMMedBench, the first large-scale Bangla biomedical Multiple Choice Question (MCQ) datasets designed to evaluate reasoning and retrieval-based methods such as Retrieval-Augmented Generation (RAG) for Bangla Question Answering. Dataset Details Dataset Description BanglaMedQA consists of 1,000 MCQs collected from authentic Bangladeshi medical admission exams (MBBS, BDS… See the full description on the dataset page: https://huggingface.co/datasets/ajwad-abrar/BanglaMedQA.text1K<n<10K0 likes110 downloads11mo agoHugging Face07mehedihasanbijoy /BanglaSEC BanglaSEC A 1.18M-pair parallel corpus for Bangla spelling error correction, with character-level error masks across 14 error types. BanglaSEC is the corpus introduced in A transformer based spelling error correction framework for Bangla and resource scarce Indic languages (Bijoy, Hossain, Islam & Shatabda, Computer Speech & Language 89:101703, 2025). Each row pairs a correct Bangla word with an erroneous form, labelled by error type and annotated with a binary mask marking… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaSEC.tabulartext-generation1M<n<10M0 likes106 downloads13d agoHugging Face08Mahadih534 /Bangladeshi-Restaurant-Datatabular10K<n<100K0 likes97 downloads2y agoHugging Face09Mahadih534 /all-bangladeshi-hospitalstabular10K<n<100K0 likes91 downloads1y agoHugging Face10Badhon /BanglaAirlineIntent Bangla / English / Banglish Airline Support Intent Classification A 19-intent classification dataset for an airline customer-support chatbot, modelled on US-Bangla Airlines' passenger mix and covering the four ways those passengers actually write: script example rows bn Bengali script আমার ফ্লাইট কি সময়মতো ছাড়বে 3,174 en English has BG147 landed yet 3,079 bl Banglish (romanized Bangla) amar flight ta ki time mto charbe 3,062 mx code-mixed mid-sentence BG435 ki… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaAirlineIntent.texttext-classification1K<n<10K1 likes82 downloads28d agoHugging Face11aplycaebous /BanglaTLit BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla Dataset Overview BanglaTLit-PT: A pre-training corpus with 245727 transliterated or romanized Bangla samples for further pre-training language models. BanglaTLit: Subset of the BanglaTLit-PT dataset containing 42705 romanized Bangla and its corresponding Bangla back-transliteration pairs. Data Description Column Title Description id A unique identifier for each data point… See the full description on the dataset page: https://huggingface.co/datasets/aplycaebous/BanglaTLit.text100K<n<1M0 likes75 downloads2y agoHugging Face12tasfuuu19 /BanglaSleep-CoT BanglaSleep-CoT The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces. Built for the Uncharted Data Challenge by Adaption Labs. Expanded using Adaptive Data by Adaption. Dataset at a Glance Why This Dataset Exists Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.tabulartext-generation1K<n<10K0 likes71 downloads5mo agoHugging Face13nahid-hub /BanglaGEC BanglaGEC: A Large-Scale Parallel Corpus for Bangla Grammatical Error Correction BanglaGEC is a large-scale parallel corpus of 7,074,425 (~7.1M) sentence pairs for Bangla (Bengali) Grammatical Error Correction (GEC). Each pair maps a grammatically erroneous Bangla sentence to its grammatically correct counterpart, along with the error type, making it directly usable for training and evaluating sequence-to-sequence models, transformers, and large language models on the Bangla… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BanglaGEC.texttext-generation1M<n<10M0 likes66 downloads2mo agoHugging Face14Mahadih534 /all-Bangladeshi-medicinestext10K<n<100K1 likes61 downloads2y agoHugging Face15mehedihasanbijoy /BanglaPRCorpus BanglaPRCorpus A 1.48M-pair corpus for Bangla punctuation restoration — unpunctuated source sentences paired with their fully punctuated targets, labelled by how many punctuation marks were removed. BanglaPRCorpus is the corpus introduced in Advancing Bangla Punctuation Restoration by a Monolingual Transformer-Based Method and a Large-Scale Corpus (Bijoy et al., EMNLP 2023 Workshop on Bangla Language Processing), alongside the Jatikarok model. Each row is a (source, target)… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaPRCorpus.texttext-generation1M<n<10M0 likes58 downloads13d agoHugging Face16Asib27 /dart_math_banglaThe dataset contains math problems in bangla. hkust-nlp/dart-math-uniform is translated using facebook/nllb-200-3.3B. To achive better performance english sentences are splitted and then fed into the translation model. textquestion-answering1K<n<10K0 likes57 downloads2y agoHugging Face17mdsajjadullah /bangla-phishing-detection-2026 Bangla Phishing Detection Dataset (SMS, Email, URLs) 2026 Synthetic dataset (~2000 rows) of phishing and legitimate messages in Bangla (Bengali) + some English, simulating common Bangladesh scams (bKash, Nagad, Daraz, Eid offers, job fraud, account lock alerts, etc.). Research Motivation Phishing/smishing attacks are rising in Bangladesh and South Asia, often in Bangla using local services. Most phishing datasets are English-only and miss these patterns.This dataset fills… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/bangla-phishing-detection-2026.tabulartext-classification1K<n<10K0 likes43 downloads7mo agoHugging Face18ahnafch01 /Bangladesh_Locations Bangladesh Postcodes Dataset (Bilingual & Structured) A comprehensive, cleaned, and bilingual (English & Bangla) database of postal codes in Bangladesh. This dataset covers the full administrative hierarchy: Division > District > Thana (Upazila) > Post Office. 📂 Files Included Filename Format Description bangladesh_postcodes_final.csv CSV The master dataset with all columns. Best for data analysis or database imports. bangladesh_postcodes_flat.json JSON A… See the full description on the dataset page: https://huggingface.co/datasets/ahnafch01/Bangladesh_Locations.tabular1K<n<10K0 likes41 downloads10mo agoHugging Face19fahim-ling /Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLPgated Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.texttranslationn<1K1 likes41 downloads25d agoHugging Face20jonybepary /Bangladesh-Voter-Synthetic-Dataset 🗳️ Bangladesh Voter Dataset 📜 Dataset Description The Bangladesh Voter Dataset is a synthetic dataset containing voter information for the purpose of demonstrating data generation and processing techniques. Each voter record includes both Bengali and English names, gender, NID, address, and profile information. 📊 Dataset Structure The dataset is structured as follows: profile: A URL to the voter's profile image. nid: A unique National Identification Number.… See the full description on the dataset page: https://huggingface.co/datasets/jonybepary/Bangladesh-Voter-Synthetic-Dataset.tabular100K<n<1M0 likes31 downloads2y agoHugging Face21faisal4590aziz /bangla-health-related-paraphrased-dataset Dataset Card for "BanglaHealthParaphrase" BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.tabulartext-generation100K<n<1M2 likes30 downloads1y agoHugging Face22shetumohanto /doctor_qa_banglatext1K<n<10K2 likes30 downloads2y agoHugging Face23kawchar85 /Bangla-Math Overview The Bangla-Math Dataset is a valuable resource that addresses the critical need for Bangla-language mathematical problem-solving datasets. Currently, there are no publicly available datasets for math problems in Bangla, making this dataset a unique and valuable contribution to the fields of natural language processing (NLP). This dataset bridges the gap by enabling AI models to understand and work effectively with Bengali math content, thereby supporting advancements in… See the full description on the dataset page: https://huggingface.co/datasets/kawchar85/Bangla-Math.tabular1K<n<10K9 likes30 downloads1d agoHugging Face24mHossain /BanglaBooktabular100K<n<1M2 likes29 downloads3y agoHugging Face25bipinsaha /bangla-law-qnatextn<1K2 likes29 downloads2y agoHugging Face26MBMMurad /Bangla_Person_Name_Extractortext1K<n<10K0 likes28 downloads3y agoHugging Face27kit-nlp /Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_BengaliA list of Chittagonian Dialect of Bangla vulgar words If you use Vulgar Lexicon dataset, please cite the following paper: @Article{app132111875, AUTHOR = {Mahmud, Tanjim and Ptaszynski, Michal and Masui, Fumito}, TITLE = {Automatic Vulgar Word Extraction Method with Application to Vulgar Remark Detection in Chittagonian Dialect of Bangla}, JOURNAL = {Applied Sciences}, VOLUME = {13}, YEAR = {2023}, NUMBER = {21}, ARTICLE-NUMBER = {11875}, URL = {https://www.mdpi.com/2076-3417/13/21/11875}… See the full description on the dataset page: https://huggingface.co/datasets/kit-nlp/Vulgar_Lexicon_of_Chittagonian_Dialect_of_Bangla_or_Bengali.texttext-classification1K<n<10K1 likes28 downloads3y agoHugging Face28musfiqdehan /preprocessed-BanglaNMTtext1M<n<10M0 likes26 downloads3y agoHugging Face29khondoker /reveal-bangla Reveal-Bangla: Intro Contains the Bangla translation of the subset from the reveal dataset. Please refer to the following code snippet which has been used to select the subset: SELECT * FROM eval Where ( answer_model = 'Flan-UL2-20B' or answer_model = 'GPT-3' AND answer_is_fully_attributable_and_correct = TRUE ); Only the following columns has been translated for the sake of the task: question full_answer step evidence Usage To load the dataset: ! pip… See the full description on the dataset page: https://huggingface.co/datasets/khondoker/reveal-bangla.tabulartext-classificationn<1K2 likes26 downloads2y agoHugging Face30md-nishat-008 /Bangla-Code-Instruct 🐯 Bangla-Code-Instruct: A Comprehensive Bangla Code Instruction Dataset Accepted at LREC 2026 Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri George Mason University, Fairfax, VA, USA The first large-scale Bangla code instruction dataset (300K examples) for training Code LLMs in Bangla. ⚠️ Note: The dataset will be released after the LREC 2026 conference. Stay tuned! Overview Bangla-Code-Instruct is a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Code-Instruct.texttext-generation100K<n<1M0 likes26 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.