CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BanglaLLM /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.tabulartext-generationn<1K0 likes188 downloads1mo agoHugging Face02momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads24d agoHugging Face03csebuetnlp /BanglaContextualBias Dataset Card for Bangla Contextual Bias The Bangla Contextual Bias dataset corresponds to the data described in the paper "An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla" accepted in ACL 2024 (Findings). Dataset Description The dataset has different parts for different bias detection experiments conducted for Bengali. WEAT & SEAT For the WEAT experiment, the dataset is translated from its English counterpart and… See the full description on the dataset page: https://huggingface.co/datasets/csebuetnlp/BanglaContextualBias.textsentence-similarityn<1K1 likes168 downloads2y agoHugging Face04nymtheescobar /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.tabulartext-generationn<1K0 likes142 downloads1mo agoHugging Face05Starscream-11813 /BanglaBook BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews This repository contains the code, data, and models of the paper titled "BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews" published in the Findings of the Association for Computational Linguistics: ACL 2023. License: Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Data Format Each row consists of a book review sample. The… See the full description on the dataset page: https://huggingface.co/datasets/Starscream-11813/BanglaBook.texttext-classification100K<n<1M2 likes99 downloads1y agoHugging Face06aridhasan /BanglaMultiHate BanglaMultiHate: Multi-task Bangla Hate-speech Dataset The BanglaMultiHate dataset collected public comments from YouTube videos using the YouTube API, primarily from Somoy TV, which is a popular Bangla News channel. The comments belong to 19 different categories, including Business, Celebrities, Disaster, Entertainment, Fashion, Geopolitics, Health, History, International, Lifestyle, Literature, Miscellaneous, National, Opinion, Politics, Religion, Science, Sports, and Technology… See the full description on the dataset page: https://huggingface.co/datasets/aridhasan/BanglaMultiHate.texttext-classification10K<n<100K0 likes99 downloads5mo agoHugging Face07kishormorol /bangla-nlp-catalog Bangla NLP Catalog A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link. This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them. Why this exists Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.tabulartext-classificationn<1K0 likes85 downloads10d agoHugging Face08zmsali /bangla-dialect-normalization Bangla Dialect Normalization Dataset A parallel corpus mapping standard Bangla to five regional Bangla dialects, built from the Vashantor dataset. Each row contains the same sentence in standard Bangla and Banglish (romanized), alongside its dialect Bangla and dialect Banglish equivalent, plus an English gloss. Regions covered Barishal, Chittagong, Mymensingh, Noakhali, Sylhet Schema Field Description standard_bangla Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.texttranslation10K<n<100K0 likes74 downloads26d agoHugging Face09abubakar-siddik /bangla-alpaca Bangla Alpaca Bangla Alpaca is a culturally localized Bangla (বাংলা) adaptation of the Stanford Alpaca dataset. Unlike simple translation, this dataset uses native-first localization to produce natural, conversational Bangla instruction-following data for training high-quality LLMs. 📊 Overview Aspect Description Language Bangla (বাংলা) Format Instruction-Input-Output Samples ~52K License Apache 2.0 📁 Dataset Structure {… See the full description on the dataset page: https://huggingface.co/datasets/abubakar-siddik/bangla-alpaca.texttext-generation10K<n<100K0 likes60 downloads9mo agoHugging Face103amthoughts /hsc-zoology-bangla-comprehensive-dataset 🧬 HSC Zoology Bangla Comprehensive Dataset A Diverse Multi-Chapter Academic Dataset This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum. 📚 Chapters Covered Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.textquestion-answering10K<n<100K1 likes54 downloads3mo agoHugging Face11shuvo-xyz /BanglaCEH BanglaCEH: A Benchmark for Culturally Entangled Homograph Disambiguation in Bangla BanglaCEH is the benchmark released with the paper "When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs." Many Bangla words are simultaneously a personal name and a culturally loaded common noun. মায়া (Maya) is both a common girl's name and a word for deep affectionate compassion; আরিফ (Arif) is a boy's name and… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-xyz/BanglaCEH.texttoken-classification1K<n<10K1 likes43 downloads1mo agoHugging Face12reyazul /BanglaSTEMtext1K<n<10K1 likes40 downloads11mo agoHugging Face13istiaqfuad /bangla-english-banglish-pairs Bangla-English-Banglish Trilingual Pairs Overview This dataset provides contrastive training pairs for fine-tuning trilingual (Bangla / Banglish / English) sentence embedding models (such as BGE-M3). It is designed to impart robustness to Banglish spelling variation. The dataset is combined from two main sources: LLM-generated Banglish spelling variants. The OPUS-100 EN-BN parallel corpus. Included Files File Rows Size Description… See the full description on the dataset page: https://huggingface.co/datasets/istiaqfuad/bangla-english-banglish-pairs.textfeature-extraction1M<n<10M0 likes39 downloads4mo agoHugging Face14subhajitmahata84 /banglabridge-instructions Dataset Card — BanglaBridge Banglish Instruction Set Summary An original instruction-tuning dataset for code-mixed / romanized Bengali ("Banglish") — the register 100M+ people actually type online (e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by safe, deterministic transformation of our own templates. Nothing is scraped, so the whole set is free to redistribute on Hugging Face and Kaggle. This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.texttext-generationn<1K0 likes39 downloads3mo agoHugging Face15tanim494 /BangladeshiVQA BangladeshiVQA A culturally grounded Bangla Visual Question Answering benchmark. BangladeshiVQA is a native Bangla VQA benchmark of 2,068 Bangladeshi images and 7,038 open-ended question–answer pairs, organized into three cognitive levels and seven image categories. To our knowledge it is the first Bangla VQA dataset with dedicated in-image Bangla scene-text (OCR) questions, and the first to split strictly by image ID to prevent train/test leakage. This Hugging Face repository… See the full description on the dataset page: https://huggingface.co/datasets/tanim494/BangladeshiVQA.imagevisual-question-answering1K<n<10K0 likes37 downloads1mo agoHugging Face163amthoughts /hsc-biology-bangla-dataset 🌿 HSC Biology Bangla Dataset (Plant Physiology) The Ultimate Resource for Bengali STEM NLP This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants. ✨ Key Highlights Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.textquestion-answering10K<n<100K1 likes34 downloads4mo agoHugging Face17Sadatsami /bangladesh-law-professional 🇧🇩 Bangladesh Law Professional Dataset A clean, instruction-tuned (Alpaca-style) question–answer dataset for fine-tuning language models on Bangladesh law, in Bangla and English. 👤 Author & Contribution Curated & built by Sadat Sami (@Sadatsami) Role Dataset architect — collected, cleaned, filtered, reformatted and published Motivation Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.textquestion-answering1K<n<10K0 likes33 downloads2mo agoHugging Face18md-nishat-008 /MBPP-Bangla 🐯 MBPP-Bangla: A Benchmark for Evaluating Bangla Code Generation Accepted at LREC 2026 Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri George Mason University, Fairfax, VA, USA The first expert-validated, multi-language Bangla code generation benchmark with 974 problems across 5 programming languages. ⚠️ Note: The benchmark will be released after the LREC 2026 conference. Stay tuned! Overview MBPP-Bangla is a… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/MBPP-Bangla.texttext-generationn<1K0 likes31 downloads6mo agoHugging Face19nihalbaig /alpaca_banglatext10K<n<100K2 likes26 downloads3y agoHugging Face20sayurio /bangla-wikipedia Bangla (Bengali) Wikipedia Articles Dataset Request More ScrapesOrder Private Scrapes Current Progress: Approx 20% Dataset Summary This dataset contains a comprehensive extraction of articles from the Bangla (Bengali) Wikipedia. It is designed for Natural Language Processing (NLP) tasks, linguistic research, and training Large Language Models (LLMs) to better understand and generate the Bengali language. Copyright and Fair Use I do not own the… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-wikipedia.texttext-generation100K<n<1M2 likes25 downloads6mo agoHugging Face21millat /indian_university_guidance_for_bangladeshi_students Indian University Guidance for Bangladeshi Students Dataset Dataset Description This dataset contains 7,044 high-quality, instruction-formatted Question-Answer pairs designed for fine-tuning Large Language Models (LLMs). The primary goal of this dataset is to create a specialized AI counselor that provides accurate, culturally relevant, and comprehensive guidance on Indian universities for Bangladeshi students. The dataset was generated through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/millat/indian_university_guidance_for_bangladeshi_students.text1K<n<10K1 likes24 downloads1y agoHugging Face22nahidstaq /bangla-llm-datagated Bangla NLP Text Corpus — 800K+ Bangla Text Samples for LLM Training and NLP Research The largest open, multi-domain Bangla text dataset, combining 801,645 samples from 15 different sources — newspapers, social media, education, reviews, QA, medical, poetry, and more. Ready for Bangla LLM pretraining, fine-tuning, and downstream NLP tasks. Why This Dataset? Bangla (Bengali) is the 7th most spoken language in the world with 230M+ speakers, but high-quality Bangla NLP… See the full description on the dataset page: https://huggingface.co/datasets/nahidstaq/bangla-llm-data.text100K<n<1M2 likes23 downloads6mo agoHugging Face23Roy2022331060 /BanglaMultiHate BanglaMultiHate: Multi-task Bangla Hate-speech Dataset The BanglaMultiHate dataset collected public comments from YouTube videos using the YouTube API, primarily from Somoy TV, which is a popular Bangla News channel. The comments belong to 19 different categories, including Business, Celebrities, Disaster, Entertainment, Fashion, Geopolitics, Health, History, International, Lifestyle, Literature, Miscellaneous, National, Opinion, Politics, Religion, Science, Sports, and… See the full description on the dataset page: https://huggingface.co/datasets/Roy2022331060/BanglaMultiHate.texttext-classification10K<n<100K0 likes23 downloads1mo agoHugging Face24rasheduzzaman /Bangla_question_answer_pair_70K_datasettext10K<n<100K1 likes22 downloads2y agoHugging Face25Badhon /BanglaQuranPunctuationDataset Bangla Quran Punctuation Dataset A high-quality dataset for Bangla punctuation restoration, derived exclusively from the Bangla translation of the Holy Quran. Dataset Description This dataset is designed for training models on punctuation restoration in Bangla text. Each sample consists of: human: Bangla text with all punctuation removed gpt: The original Bangla text with correct punctuation (। , ; : - ! ?) All samples are derived from the Bangla translation of the… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaQuranPunctuationDataset.textfill-mask1K<n<10K1 likes22 downloads9mo agoHugging Face26adnan1837 /Bangla_jokesgated Dataset Card for Bangla Jokes Dataset The Bangla Jokes Dataset is a collection of humorous text samples written in Bengali (Bangla). This dataset is intended for NLP research and model training, especially in the area of Bangla-language humor generation, sentiment, or cultural studies. It is one of the first attempts to gather a sizable dataset of jokes in Bangla for open-source use. Dataset Details Curated by: Adnan1837 Funded by: No one Shared by: Adnan, Md.… See the full description on the dataset page: https://huggingface.co/datasets/adnan1837/Bangla_jokes.textn<1K1 likes20 downloads1y agoHugging Face27azminetoushikwasi /bangla-bcs-qstexttext-classification1K<n<10K0 likes18 downloads2y agoHugging Face28Tensoic /GPTeacher-Banglatext10K<n<100K0 likes17 downloads3y agoHugging Face29Badhon /BanglaPunctDataset Bangla Punctuation Restoration Dataset A merged, high-quality Bangla dataset for punctuation restoration, formatted as instruction-tuning conversation pairs.The dataset is suitable for fine-tuning Large Language Models (LLMs) and sequence models to restore punctuation in Bangla text. Dataset Summary Language: Bengali (Bangla) Task: Punctuation Restoration Format: JSONL (instruction-style conversations) Max chunk length: ~256 characters Punctuation covered:। ! ? , ; : -… See the full description on the dataset page: https://huggingface.co/datasets/Badhon/BanglaPunctDataset.texttoken-classification10K<n<100K0 likes17 downloads9mo agoHugging Face30kawsarahmd /xlsum_bangla Dataset Card for kawsarahmd/papers_summary_datasets_xsum_bangla This dataset is derived from csebuetnlp/xlsum (subset: bengali). Dataset Description Overview Original dataset: csebuetnlp/xlsum Subset: bengali Total samples: 10126 Split source: Original splits from dataset Splits Train split: 8102 samples (80.0%) Validation split: 1012 samples (10.0%) Test split: 1012 samples (10.0%) Features { "id": "object", "url": "object"… See the full description on the dataset page: https://huggingface.co/datasets/kawsarahmd/xlsum_bangla.text10K<n<100K0 likes15 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.