CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hishab /titulm-bangla-corpus TituLM Bangla Corpus This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability. This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/hishab/titulm-bangla-corpus.texttext-generation10M<n<100M13 likes1.8k downloads1y agoHugging Face02Eamin-sust /BanglaEng-SynCorpus BanglaEng-SynCorpus Dataset Summary BanglaEng-SynCorpus is a large-scale synthetic Bangla–English parallel corpus designed to support research in Neural Machine Translation (NMT) and other Bangla–English bilingual NLP tasks.The corpus is generated using linguistically validated sentence templates combined with topic-wise curated vocabularies, covering all 12 English/Bangla tense structures. Due to extreme scale (trillions of possible sentence pairs), the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Eamin-sust/BanglaEng-SynCorpus.texttranslation10B<n<100B1 likes1.5k downloads9mo agoHugging Face03shofikul-1234 /titulm-bangla-corpus TituLM Bangla Corpus This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability. This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we… See the full description on the dataset page: https://huggingface.co/datasets/shofikul-1234/titulm-bangla-corpus.texttext-generation10M<n<100M0 likes909 downloads3mo agoHugging Face04kamruzzaman-asif /bangla-instruction-dataset 🧠 Bangla Instruction Dataset This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models. 📚 Dataset Splits The dataset is organized into the following splits: Split Name Source Dataset Description OdiaGenAI OdiaGenAI/all_combined_bengali_252k A large-scale collection of diverse Bangla instructions and responses. chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.texttext-generation1M<n<10M1 likes210 downloads1y agoHugging Face05FaiyazAbdullah114708 /BanglaVerse Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision–language… See the full description on the dataset page: https://huggingface.co/datasets/FaiyazAbdullah114708/BanglaVerse.imagetranslation10K<n<100K2 likes198 downloads6mo agoHugging Face06BanglaLLM /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.tabulartext-generationn<1K0 likes188 downloads1mo agoHugging Face07momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads24d agoHugging Face08md-nishat-008 /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.texttext-generation10K<n<100K2 likes148 downloads1y agoHugging Face09nymtheescobar /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.tabulartext-generationn<1K0 likes142 downloads1mo agoHugging Face10munzurul /bangla-corpus TituLM Bangla Corpus This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability. This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/munzurul/bangla-corpus.texttext-generation10M<n<100M0 likes141 downloads7mo agoHugging Face11md-nishat-008 /Bangla-Instruct Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.texttext-generation100K<n<1M8 likes119 downloads1y agoHugging Face12Mahadih534 /Institutional-Information-of-Bangladesh Institutional-Information-of-Bangladesh Dataset This Dataset contains all verified and authorized Institutional information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.tabularquestion-answering10K<n<100K2 likes114 downloads2y agoHugging Face13mehedihasanbijoy /BanglaSEC BanglaSEC A 1.18M-pair parallel corpus for Bangla spelling error correction, with character-level error masks across 14 error types. BanglaSEC is the corpus introduced in A transformer based spelling error correction framework for Bangla and resource scarce Indic languages (Bijoy, Hossain, Islam & Shatabda, Computer Speech & Language 89:101703, 2025). Each row pairs a correct Bangla word with an erroneous form, labelled by error type and annotated with a binary mask marking… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaSEC.tabulartext-generation1M<n<10M0 likes106 downloads13d agoHugging Face14Hasin2026 /bangladesh-scob-judgment-summarization Bangladesh Supreme Court (SCOB) High Court Division Judgment Summarization Dataset Dataset Summary The Bangladesh Supreme Court (SCOB) High Court Division Judgment Summarization Dataset is a curated, high-quality legal NLP dataset comprising all 235 canonical judgments published in the Supreme Court Online Bulletin (SCOB) by the High Court Division of the Supreme Court of Bangladesh. Each sample pairs a complete, cleaned legal judgment body with its official… See the full description on the dataset page: https://huggingface.co/datasets/Hasin2026/bangladesh-scob-judgment-summarization.textsummarizationn<1K0 likes78 downloads19d agoHugging Face15tasfuuu19 /BanglaSleep-CoT BanglaSleep-CoT The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces. Built for the Uncharted Data Challenge by Adaption Labs. Expanded using Adaptive Data by Adaption. Dataset at a Glance Why This Dataset Exists Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.tabulartext-generation1K<n<10K0 likes71 downloads5mo agoHugging Face16ihumaunkabir /alpaca-gpt4-bangla alpaca-gpt4-bangla A Bangla (Bengali) instruction-following dataset for supervised fine-tuning (SFT) of large language models. It contains ~49,969 instruction-response pairs covering a wide range of topics -- coding, creative writing, reasoning, summarization, math, open Q&A -- suitable for teaching a base model to follow instructions in Bangla. This dataset is a Korean -> Bangla machine translation of FreedomIntelligence/alpaca-gpt4-korean, which is itself a Korean translation… See the full description on the dataset page: https://huggingface.co/datasets/ihumaunkabir/alpaca-gpt4-bangla.texttext-generation10K<n<100K0 likes68 downloads3mo agoHugging Face17nahid-hub /BanglaGEC BanglaGEC: A Large-Scale Parallel Corpus for Bangla Grammatical Error Correction BanglaGEC is a large-scale parallel corpus of 7,074,425 (~7.1M) sentence pairs for Bangla (Bengali) Grammatical Error Correction (GEC). Each pair maps a grammatically erroneous Bangla sentence to its grammatically correct counterpart, along with the error type, making it directly usable for training and evaluating sequence-to-sequence models, transformers, and large language models on the Bangla… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BanglaGEC.texttext-generation1M<n<10M0 likes66 downloads2mo agoHugging Face18abubakar-siddik /bangla-alpaca Bangla Alpaca Bangla Alpaca is a culturally localized Bangla (বাংলা) adaptation of the Stanford Alpaca dataset. Unlike simple translation, this dataset uses native-first localization to produce natural, conversational Bangla instruction-following data for training high-quality LLMs. 📊 Overview Aspect Description Language Bangla (বাংলা) Format Instruction-Input-Output Samples ~52K License Apache 2.0 📁 Dataset Structure {… See the full description on the dataset page: https://huggingface.co/datasets/abubakar-siddik/bangla-alpaca.texttext-generation10K<n<100K0 likes60 downloads9mo agoHugging Face19mehedihasanbijoy /BanglaPRCorpus BanglaPRCorpus A 1.48M-pair corpus for Bangla punctuation restoration — unpunctuated source sentences paired with their fully punctuated targets, labelled by how many punctuation marks were removed. BanglaPRCorpus is the corpus introduced in Advancing Bangla Punctuation Restoration by a Monolingual Transformer-Based Method and a Large-Scale Corpus (Bijoy et al., EMNLP 2023 Workshop on Bangla Language Processing), alongside the Jatikarok model. Each row is a (source, target)… See the full description on the dataset page: https://huggingface.co/datasets/mehedihasanbijoy/BanglaPRCorpus.texttext-generation1M<n<10M0 likes58 downloads13d agoHugging Face203amthoughts /hsc-zoology-bangla-comprehensive-dataset 🧬 HSC Zoology Bangla Comprehensive Dataset A Diverse Multi-Chapter Academic Dataset This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum. 📚 Chapters Covered Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.textquestion-answering10K<n<100K1 likes54 downloads3mo agoHugging Face21sayurio /ekpatagolpo-scrape-bangla-literature Ekpatagolpo Bengali Stories Archive Request More ScrapesOrder Private Scrapes Overview This repository contains a large-scale, curated text dataset scraped from ekpatagolpo.com. The primary goal of this archive is to preserve a massive collection of purely human-written Bengali literature and stories (Bangla Golpo), creating a distinct record of human creativity separate from AI-generated text. Purpose and Usage This dataset is published publicly under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/ekpatagolpo-scrape-bangla-literature.texttext-generation10K<n<100K1 likes50 downloads6mo agoHugging Face22shuvo-xyz /BanglaCEH BanglaCEH: A Benchmark for Culturally Entangled Homograph Disambiguation in Bangla BanglaCEH is the benchmark released with the paper "When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs." Many Bangla words are simultaneously a personal name and a culturally loaded common noun. মায়া (Maya) is both a common girl's name and a word for deep affectionate compassion; আরিফ (Arif) is a boy's name and… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-xyz/BanglaCEH.texttoken-classification1K<n<10K1 likes43 downloads1mo agoHugging Face23subhajitmahata84 /banglabridge-instructions Dataset Card — BanglaBridge Banglish Instruction Set Summary An original instruction-tuning dataset for code-mixed / romanized Bengali ("Banglish") — the register 100M+ people actually type online (e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by safe, deterministic transformation of our own templates. Nothing is scraped, so the whole set is free to redistribute on Hugging Face and Kaggle. This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.texttext-generationn<1K0 likes39 downloads3mo agoHugging Face24sayurio /jugantor.com-scrape-bangla Jugantor News Archive (Bangla) Overview This repository contains a comprehensive text dataset scraped from jugantor.com, one of the leading Bengali daily newspapers in Bangladesh. The primary goal of this archive is to preserve a massive collection of purely human-written journalism, editorials, and news reports, creating a distinct record of human-authored text separate from AI-generated content. Purpose and Usage This dataset is published publicly and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/jugantor.com-scrape-bangla.imagetext-generation10K<n<100K1 likes36 downloads6mo agoHugging Face25KillerShoaib /DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-DataDeepSeek R1 Bangla MMLU Distil Dataset Original Dataset: hishab/bangla-mmlu Train Samples: 17,796 Test Samples: 2,576 Total API Cost: 7K BDT Contributors: Myself Numaer How the Dataset was created Step 1 - Base Dataset I've used bangla-mmlu dataset released by hisab. Kudos to them for creating and open sourcing the dataset. Without their dataset this synthetic reasoning dataset won't exist in the first place. Step 2 - Select Subset Since I'm… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-Data.textquestion-answering10K<n<100K14 likes35 downloads1y agoHugging Face263amthoughts /hsc-biology-bangla-dataset 🌿 HSC Biology Bangla Dataset (Plant Physiology) The Ultimate Resource for Bengali STEM NLP This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants. ✨ Key Highlights Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.textquestion-answering10K<n<100K1 likes34 downloads4mo agoHugging Face27Sadatsami /bangladesh-law-professional 🇧🇩 Bangladesh Law Professional Dataset A clean, instruction-tuned (Alpaca-style) question–answer dataset for fine-tuning language models on Bangladesh law, in Bangla and English. 👤 Author & Contribution Curated & built by Sadat Sami (@Sadatsami) Role Dataset architect — collected, cleaned, filtered, reformatted and published Motivation Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.textquestion-answering1K<n<10K0 likes33 downloads2mo agoHugging Face28md-nishat-008 /MBPP-Bangla 🐯 MBPP-Bangla: A Benchmark for Evaluating Bangla Code Generation Accepted at LREC 2026 Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri George Mason University, Fairfax, VA, USA The first expert-validated, multi-language Bangla code generation benchmark with 974 problems across 5 programming languages. ⚠️ Note: The benchmark will be released after the LREC 2026 conference. Stay tuned! Overview MBPP-Bangla is a… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/MBPP-Bangla.texttext-generationn<1K0 likes31 downloads6mo agoHugging Face29faisal4590aziz /bangla-health-related-paraphrased-dataset Dataset Card for "BanglaHealthParaphrase" BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.tabulartext-generation100K<n<1M2 likes30 downloads1y agoHugging Face30spitfire4794 /Bangla-SFT-50k Bangla-SFT Bangla-SFT is an instruction-following dataset containing 50,053 Bengali prompt-response pairs. It was scaled up from a 500-sample seed dataset (spitfire4794/bang_seed). Dataset Summary The dataset covers 6 task categories. The prompts are designed to be self-contained (hydrated with appropriate contextual inputs), and the responses are formatted to be direct, omitting conversational prefaces and filler. Seed Generation: Baseline instructions generated… See the full description on the dataset page: https://huggingface.co/datasets/spitfire4794/Bangla-SFT-50k.texttext-generation10K<n<100K1 likes28 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.