CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BanglaLLM /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.tabulartext-generationn<1K0 likes188 downloads1mo agoHugging Face02momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads24d agoHugging Face03nymtheescobar /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.tabulartext-generationn<1K0 likes142 downloads1mo agoHugging Face04abubakar-siddik /bangla-alpaca Bangla Alpaca Bangla Alpaca is a culturally localized Bangla (বাংলা) adaptation of the Stanford Alpaca dataset. Unlike simple translation, this dataset uses native-first localization to produce natural, conversational Bangla instruction-following data for training high-quality LLMs. 📊 Overview Aspect Description Language Bangla (বাংলা) Format Instruction-Input-Output Samples ~52K License Apache 2.0 📁 Dataset Structure {… See the full description on the dataset page: https://huggingface.co/datasets/abubakar-siddik/bangla-alpaca.texttext-generation10K<n<100K0 likes60 downloads9mo agoHugging Face053amthoughts /hsc-zoology-bangla-comprehensive-dataset 🧬 HSC Zoology Bangla Comprehensive Dataset A Diverse Multi-Chapter Academic Dataset This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum. 📚 Chapters Covered Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.textquestion-answering10K<n<100K1 likes54 downloads3mo agoHugging Face06shuvo-xyz /BanglaCEH BanglaCEH: A Benchmark for Culturally Entangled Homograph Disambiguation in Bangla BanglaCEH is the benchmark released with the paper "When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs." Many Bangla words are simultaneously a personal name and a culturally loaded common noun. মায়া (Maya) is both a common girl's name and a word for deep affectionate compassion; আরিফ (Arif) is a boy's name and… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-xyz/BanglaCEH.texttoken-classification1K<n<10K1 likes43 downloads1mo agoHugging Face07subhajitmahata84 /banglabridge-instructions Dataset Card — BanglaBridge Banglish Instruction Set Summary An original instruction-tuning dataset for code-mixed / romanized Bengali ("Banglish") — the register 100M+ people actually type online (e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by safe, deterministic transformation of our own templates. Nothing is scraped, so the whole set is free to redistribute on Hugging Face and Kaggle. This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.texttext-generationn<1K0 likes39 downloads3mo agoHugging Face083amthoughts /hsc-biology-bangla-dataset 🌿 HSC Biology Bangla Dataset (Plant Physiology) The Ultimate Resource for Bengali STEM NLP This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants. ✨ Key Highlights Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.textquestion-answering10K<n<100K1 likes34 downloads4mo agoHugging Face09Sadatsami /bangladesh-law-professional 🇧🇩 Bangladesh Law Professional Dataset A clean, instruction-tuned (Alpaca-style) question–answer dataset for fine-tuning language models on Bangladesh law, in Bangla and English. 👤 Author & Contribution Curated & built by Sadat Sami (@Sadatsami) Role Dataset architect — collected, cleaned, filtered, reformatted and published Motivation Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.textquestion-answering1K<n<10K0 likes33 downloads2mo agoHugging Face10md-nishat-008 /MBPP-Bangla 🐯 MBPP-Bangla: A Benchmark for Evaluating Bangla Code Generation Accepted at LREC 2026 Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri George Mason University, Fairfax, VA, USA The first expert-validated, multi-language Bangla code generation benchmark with 974 problems across 5 programming languages. ⚠️ Note: The benchmark will be released after the LREC 2026 conference. Stay tuned! Overview MBPP-Bangla is a… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/MBPP-Bangla.texttext-generationn<1K0 likes31 downloads6mo agoHugging Face11sayurio /bangla-wikipedia Bangla (Bengali) Wikipedia Articles Dataset Request More ScrapesOrder Private Scrapes Current Progress: Approx 20% Dataset Summary This dataset contains a comprehensive extraction of articles from the Bangla (Bengali) Wikipedia. It is designed for Natural Language Processing (NLP) tasks, linguistic research, and training Large Language Models (LLMs) to better understand and generate the Bengali language. Copyright and Fair Use I do not own the… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-wikipedia.texttext-generation100K<n<1M2 likes25 downloads6mo agoHugging Face12likhonhfai /bangla-stories-datasetgated Bangla Multi-Task Benchmark Dataset A comprehensive multi-task dataset for evaluating and fine-tuning LLMs on Bangla. Statistics Metric Value Total Samples 1,592 Task Types 8 Training 1,273 Validation 159 Test 160 Task Types Task Count Difficulty text_generation 200 Medium title_generation 200 Easy summarization 200 Medium comprehension 200 Hard classification 200 Easy sentiment 200 Easy keywords 200 Easy… See the full description on the dataset page: https://huggingface.co/datasets/likhonhfai/bangla-stories-dataset.texttext-generation1K<n<10K1 likes13 downloads7mo agoHugging Face13sayurio /bangla-kobita-scrape-bangla-literature Bangla Kobita Poetry Archive Overview This repository contains a curated text dataset of Bengali poetry scraped from the web, primarily targeting comprehensive poetry platforms like bangla-kobita.com. The primary goal of this archive is to preserve a rich collection of purely human-written Bengali poems (Bangla Kobita), creating a distinct record of human artistic expression, emotion, and linguistic rhythm separate from AI-generated text. Purpose and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-kobita-scrape-bangla-literature.imagetext-generation100K<n<1M1 likes12 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.