CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tanziro /bangla-crime-investigation-patterns-v2 Bangla Crime Investigation Patterns V2 This dataset is an anonymized, structured extraction from 1,078 public Bangla crime-investigation video transcripts. It was built to support machine-learning and LLM research on crime-pattern information extraction, case summarization, event sequencing, entity/relationship extraction, and Bengali investigative-report analysis. The public release does not include raw full transcripts. It contains short redacted evidence snippets linked to… See the full description on the dataset page: https://huggingface.co/datasets/tanziro/bangla-crime-investigation-patterns-v2.tabulartext-classification10K<n<100K0 likes313 downloads4mo agoHugging Face02sartajekram /BanglaRQABanglaRQA is a human-annotated Bangla Question Answering (QA) dataset with diverse question-answer types.textquestion-answering10K<n<100K7 likes298 downloads3y agoHugging Face03kamruzzaman-asif /bangla-instruction-dataset 🧠 Bangla Instruction Dataset This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models. 📚 Dataset Splits The dataset is organized into the following splits: Split Name Source Dataset Description OdiaGenAI OdiaGenAI/all_combined_bengali_252k A large-scale collection of diverse Bangla instructions and responses. chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.texttext-generation1M<n<10M1 likes210 downloads1y agoHugging Face04FaiyazAbdullah114708 /BanglaVerse Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision–language… See the full description on the dataset page: https://huggingface.co/datasets/FaiyazAbdullah114708/BanglaVerse.imagetranslation10K<n<100K2 likes198 downloads6mo agoHugging Face05momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads24d agoHugging Face06Mahadih534 /Institutional-Information-of-Bangladesh Institutional-Information-of-Bangladesh Dataset This Dataset contains all verified and authorized Institutional information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.tabularquestion-answering10K<n<100K2 likes114 downloads2y agoHugging Face07shawon /bangla-math-chat bangla-math-chat A math dataset for fine-tuning LLMs to chat on math problems in Bangla. This dataset is a reformatted version of BanglaLLM/bangla_math_by_Ashrafur. The code to reformat the original dataset can be found on Github: ShawonAshraf/bangla-math-chat textquestion-answering100K<n<1M1 likes108 downloads1y agoHugging Face08hishab /bangla-mmlu Data Summary We curated multiple-choice questions from various open-source educational websites and textbooks, inspired by the original MMLU dataset (Hendrycks et al., 2020). The dataset includes multiple-choice questions from different Bangladeshi exams, such as job exams, the Bangladesh Civil Service Exam, and undergraduate admission exams. In Figure 7, we report category wise distributions. Bangla MMLU Dataset Overview The Bangla MMLU dataset consists of a total of 116,503… See the full description on the dataset page: https://huggingface.co/datasets/hishab/bangla-mmlu.textquestion-answering10K<n<100K5 likes101 downloads1y agoHugging Face09kishormorol /bangla-nlp-catalog Bangla NLP Catalog A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link. This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them. Why this exists Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.tabulartext-classificationn<1K0 likes85 downloads10d agoHugging Face10csebuetnlp /BanglaSocialBias Dataset Card for Bangla Contextual Bias The Bangla Social Bias dataset comprises of the data used in the paper titled "Social Bias in Large Language Models For Bangla: An Empirical Study on Gender and Religious Bias". Dataset Description The dataset contains different domains of data used for the experimentations mentioned in the paper. A summary of the different categories of data provided in this dataset are: the formatted raw data collected from open source for the… See the full description on the dataset page: https://huggingface.co/datasets/csebuetnlp/BanglaSocialBias.text-classification1K<n<10K2 likes80 downloads2y agoHugging Face11afzal-hosen-mandal /index-of-the-bangladesh-code Index of the Bangladesh Code A structured, machine-readable research index of Bangladesh Code legal records. Maintainer: Afzal Hosen MandalOrganization: LegalDefenseHub / Afzal & AssociatesJurisdiction: BangladeshCoverage: 1799–2026Parsed source records: 1,674 Historical periods Period Years Records British Period 1799–1947 245 Pakistan Period 1948–1971 160 Bangladesh Period 1972–2026 1269 Total 1799–2026 1,674 Important… See the full description on the dataset page: https://huggingface.co/datasets/afzal-hosen-mandal/index-of-the-bangladesh-code.text-classification1K<n<10K1 likes80 downloads1mo agoHugging Face12tasfuuu19 /BanglaSleep-CoT BanglaSleep-CoT The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces. Built for the Uncharted Data Challenge by Adaption Labs. Expanded using Adaptive Data by Adaption. Dataset at a Glance Why This Dataset Exists Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.tabulartext-generation1K<n<10K0 likes71 downloads5mo agoHugging Face13Asib27 /dart_math_banglaThe dataset contains math problems in bangla. hkust-nlp/dart-math-uniform is translated using facebook/nllb-200-3.3B. To achive better performance english sentences are splitted and then fed into the translation model. textquestion-answering1K<n<10K0 likes57 downloads2y agoHugging Face143amthoughts /hsc-zoology-bangla-comprehensive-dataset 🧬 HSC Zoology Bangla Comprehensive Dataset A Diverse Multi-Chapter Academic Dataset This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum. 📚 Chapters Covered Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.textquestion-answering10K<n<100K1 likes54 downloads3mo agoHugging Face15Armans33115 /JobCCC-Conversational-Job-Recommendation-Bangladesh JobCCC: A Conversational Code-Mixed Corpus for Job Recommendation in Bangladesh Dataset Creators Authors: Md. Arman Hossain, Mubashir Jawad, Fariha Khandakar Moon, and Sonia Binte Siraj Supervisor: Dr. Nafis Sadeq Institution: Department of Computer Science & Engineering, East West University Dataset Summary JobCCC (Conversational Code-Mixed Corpus) is a multi-turn conversational benchmark and job recommendation dataset tailored for the… See the full description on the dataset page: https://huggingface.co/datasets/Armans33115/JobCCC-Conversational-Job-Recommendation-Bangladesh.tabularquestion-answering10K<n<100K2 likes43 downloads1mo agoHugging Face16subhajitmahata84 /banglabridge-instructions Dataset Card — BanglaBridge Banglish Instruction Set Summary An original instruction-tuning dataset for code-mixed / romanized Bengali ("Banglish") — the register 100M+ people actually type online (e.g. "kal ki plan? ami free achi"). Every pair is authored by us or produced by safe, deterministic transformation of our own templates. Nothing is scraped, so the whole set is free to redistribute on Hugging Face and Kaggle. This is the originality +… See the full description on the dataset page: https://huggingface.co/datasets/subhajitmahata84/banglabridge-instructions.texttext-generationn<1K0 likes39 downloads3mo agoHugging Face17sayurio /jugantor.com-scrape-bangla Jugantor News Archive (Bangla) Overview This repository contains a comprehensive text dataset scraped from jugantor.com, one of the leading Bengali daily newspapers in Bangladesh. The primary goal of this archive is to preserve a massive collection of purely human-written journalism, editorials, and news reports, creating a distinct record of human-authored text separate from AI-generated content. Purpose and Usage This dataset is published publicly and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/jugantor.com-scrape-bangla.imagetext-generation10K<n<100K1 likes36 downloads6mo agoHugging Face18KillerShoaib /DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-DataDeepSeek R1 Bangla MMLU Distil Dataset Original Dataset: hishab/bangla-mmlu Train Samples: 17,796 Test Samples: 2,576 Total API Cost: 7K BDT Contributors: Myself Numaer How the Dataset was created Step 1 - Base Dataset I've used bangla-mmlu dataset released by hisab. Kudos to them for creating and open sourcing the dataset. Without their dataset this synthetic reasoning dataset won't exist in the first place. Step 2 - Select Subset Since I'm… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/DeepSeek-r1-Distill-Bangla-MMLU-Reasoning-Data.textquestion-answering10K<n<100K14 likes35 downloads1y agoHugging Face193amthoughts /hsc-biology-bangla-dataset 🌿 HSC Biology Bangla Dataset (Plant Physiology) The Ultimate Resource for Bengali STEM NLP This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants. ✨ Key Highlights Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.textquestion-answering10K<n<100K1 likes34 downloads4mo agoHugging Face20Sadatsami /bangladesh-law-professional 🇧🇩 Bangladesh Law Professional Dataset A clean, instruction-tuned (Alpaca-style) question–answer dataset for fine-tuning language models on Bangladesh law, in Bangla and English. 👤 Author & Contribution Curated & built by Sadat Sami (@Sadatsami) Role Dataset architect — collected, cleaned, filtered, reformatted and published Motivation Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.textquestion-answering1K<n<10K0 likes33 downloads2mo agoHugging Face21momahadi /bangladesh-bar-council-exam-dataset Bangladesh Bar Council Exam QA Dataset: 2022-2023 Bangla-English The Bangladesh Bar Council Exam QA Dataset is a 400-question bilingual Bangla-English legal question-answering benchmark compiled from the 2022 and 2023 Bangladesh Bar Council examination materials. It is intended for legal NLP evaluation, multiple-choice question answering, retrieval experiments, and Bangladesh law language-model research. This repository contains the examination benchmark only. The larger… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-bar-council-exam-dataset.question-answeringn<1K1 likes29 downloads2mo agoHugging Face22khondoker /reveal-bangla Reveal-Bangla: Intro Contains the Bangla translation of the subset from the reveal dataset. Please refer to the following code snippet which has been used to select the subset: SELECT * FROM eval Where ( answer_model = 'Flan-UL2-20B' or answer_model = 'GPT-3' AND answer_is_fully_attributable_and_correct = TRUE ); Only the following columns has been translated for the sake of the task: question full_answer step evidence Usage To load the dataset: ! pip… See the full description on the dataset page: https://huggingface.co/datasets/khondoker/reveal-bangla.tabulartext-classificationn<1K2 likes26 downloads2y agoHugging Face23sayurio /bangla-nsfw-stories-scrape Bangla Erotic Story Scraped Dataset A comprehensive text dataset containing an archive of Bengali 18+ literature and stories scraped from various online sources. 📌 Dataset Summary This repository serves as a text dataset aggregating Bengali adult fiction. The data has been collected to preserve these stories and can be used for natural language processing (NLP) tasks, linguistic analysis of colloquial Bengali, generative text modeling, or archival purposes. ⚠️… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-nsfw-stories-scrape.text-classification1 likes26 downloads6mo agoHugging Face24Simanto123 /BanglaRQABanglaRQA is a human-annotated Bangla Question Answering (QA) dataset with diverse question-answer types.question-answering10K<n<100K0 likes26 downloads6mo agoHugging Face25Mahadih534 /Bangladeshi_National_E-Services Bangladeshi_National_E-Services Dataset This Dataset contains all verified and authorized National e-services information of Bangladeshi Government Description I have collected these all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Bangladeshi_National_E-Services.question-answering1K<n<10K0 likes25 downloads2y agoHugging Face26sijantanvir /BanglaSocialBench BanglaSocialBench Paper:BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction Links 📄 Paper: https://aclanthology.org/2026.acl-srw.22/ 💻 Code: https://github.com/sijantanvir/bangla-social-bench Citation @inproceedings{sijan-etal-2026-banglasocialbench, title = {BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social… See the full description on the dataset page: https://huggingface.co/datasets/sijantanvir/BanglaSocialBench.tabularmultiple-choice1K<n<10K0 likes24 downloads2mo agoHugging Face27Mahadih534 /Bangladeshi_Doctor_List Bangladeshi_Doctor_List Dataset This Dataset contains all verified and authorized Docto information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact mahadise01@gmail.com Linkdin:… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Bangladeshi_Doctor_List.tabularquestion-answeringn<1K1 likes18 downloads2y agoHugging Face28azminetoushikwasi /bangla-bcs-qstexttext-classification1K<n<10K0 likes18 downloads2y agoHugging Face29SrejonAhamed /BanglaBorotexttext-classification100K<n<1M0 likes18 downloads2y agoHugging Face30techoptions /Bangla_Education_Specialist_v1_11K Bangla Education Specialist v1 — 11K Dataset Description A Bangla education-domain instruction-following dataset containing 11,000+ samples designed for fine-tuning Large Language Models (LLMs) on Bengali educational question-answering tasks. Curated and processed by TechOptions. Dataset Details Property Value Language Bengali (bn) Domain Education Total Samples ~11,000 File Size ~6.48 MB Format JSONL → Parquet License Apache… See the full description on the dataset page: https://huggingface.co/datasets/techoptions/Bangla_Education_Specialist_v1_11K.texttext-generation10K<n<100K0 likes17 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.