CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hishab /titulm-bangla-corpus TituLM Bangla Corpus This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability. This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/hishab/titulm-bangla-corpus.texttext-generation10M<n<100M13 likes1.8k downloads1y agoHugging Face02Eamin-sust /BanglaEng-SynCorpus BanglaEng-SynCorpus Dataset Summary BanglaEng-SynCorpus is a large-scale synthetic Bangla–English parallel corpus designed to support research in Neural Machine Translation (NMT) and other Bangla–English bilingual NLP tasks.The corpus is generated using linguistically validated sentence templates combined with topic-wise curated vocabularies, covering all 12 English/Bangla tense structures. Due to extreme scale (trillions of possible sentence pairs), the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Eamin-sust/BanglaEng-SynCorpus.texttranslation10B<n<100B1 likes1.5k downloads9mo agoHugging Face03kawsersikder /bangladesh-stock-market-dataset Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks Author: Kawser Sikder Overview A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market. Metric Value Total Stocks 441 Total Sectors 23 Total Trading Records 1,507,388 Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.tabulartime-series-forecasting1M<n<10M1 likes1.1k downloads1mo agoHugging Face04shofikul-1234 /titulm-bangla-corpus TituLM Bangla Corpus This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability. This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we… See the full description on the dataset page: https://huggingface.co/datasets/shofikul-1234/titulm-bangla-corpus.texttext-generation10M<n<100M0 likes909 downloads3mo agoHugging Face05hishab /titulm-bangla-mmlu Titulm Bangla MMLU Read the paper for details: https://arxiv.org/abs/2502.11187 Citation @misc{nahin2025titullmsfamilybanglallms, title={TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking}, author={Shahriar Kabir Nahin and Rabindra Nath Nandi and Sagor Sarker and Quazi Sarwar Muhtaseem and Md Kowsher and Apu Chandraw Shill and Md Ibrahim and Mehadi Hasan Menon and Tareq Al Muntasir and Firoj Alam}, year={2025}, eprint={2502.11187}… See the full description on the dataset page: https://huggingface.co/datasets/hishab/titulm-bangla-mmlu.text100K<n<1M6 likes883 downloads1y agoHugging Face06SKNahin /BanglaQwen-Train-Corpustext10M<n<100M0 likes730 downloads2y agoHugging Face07Banglabox /bangla-corpus BanglaBox — Bangladeshi Bangla TTS corpus Anonymous artifact for double-blind review. A Bangladeshi Bangla speech corpus for text-to-speech and zero-shot voice cloning, built with the coverage-driven script pipeline described in the paper (7 domains — news, customer care, teaching, healthcare, e-commerce, finance, IT — with scripts selected under a tiered Jensen–Shannon-divergence objective over phones, diphones, triphones and conjunct clusters (juktakkhor) and filtered by… See the full description on the dataset page: https://huggingface.co/datasets/Banglabox/bangla-corpus.audio100K<n<1M0 likes723 downloads5d agoHugging Face08JabaleNurAdnan /bangla-noise-robustness-datatext100K<n<1M0 likes614 downloads12d agoHugging Face09vaishali /banglaTabQA Dataset Card for "banglaTabQA" Usage import pandas as pd from datasets import load_dataset banglatableQA = load_dataset("vaishali/banglaTabQA") for sample in banglatableQA['train']: question = sample['question'] input_table = pd.read_json(sample['table'], orient='split') answer = pd.read_json(sample['answer'], orient='split') BibTeX entry and citation info @inproceedings{pal-etal-2024-table, title = "Table Question Answering for Low-resourced… See the full description on the dataset page: https://huggingface.co/datasets/vaishali/banglaTabQA.texttable-question-answering1M<n<10M0 likes355 downloads2y agoHugging Face10psdn-ai /bangla-10kgated Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh Bangla-10K is a 10,816-hour Bengali speech corpus with 624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus (567,323 recordings) and a separately collected 745.1-hour evaluation set (57,628 recordings). It combines scripted single-speaker read speech with natural multi-speaker conversations for Bengali automatic speech recognition (ASR). The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.audioautomatic-speech-recognition100K<n<1M0 likes342 downloads2d agoHugging Face11justicedao /ipfs_bangladesh_laws_ir Bangladesh legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_bangladesh_laws (revision 16782096c126f7342b3cfeaa312c437c9fa2de73) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Bangladesh prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_bangladesh_laws_ir.tabulartext-retrieval100K<n<1M0 likes338 downloads14h agoHugging Face12tanziro /bangla-crime-investigation-patterns-v2 Bangla Crime Investigation Patterns V2 This dataset is an anonymized, structured extraction from 1,078 public Bangla crime-investigation video transcripts. It was built to support machine-learning and LLM research on crime-pattern information extraction, case summarization, event sequencing, entity/relationship extraction, and Bengali investigative-report analysis. The public release does not include raw full transcripts. It contains short redacted evidence snippets linked to… See the full description on the dataset page: https://huggingface.co/datasets/tanziro/bangla-crime-investigation-patterns-v2.tabulartext-classification10K<n<100K0 likes313 downloads4mo agoHugging Face13sartajekram /BanglaRQABanglaRQA is a human-annotated Bangla Question Answering (QA) dataset with diverse question-answer types.textquestion-answering10K<n<100K7 likes298 downloads3y agoHugging Face14swadhinbiswas /bangladeshi-jobs Bangladeshi Tech Jobs — Open Dataset Weekly-refreshed, structured dataset of open software & IT job postings from Bangladeshi tech companies, built by an automated crawl → LLM-extraction → data-warehouse pipeline. Published as JSON + Parquet + a DuckDB star schema, free for any use with attribution (CC-BY-4.0). Snapshot (2026-09-20) 🤖 Auto-generated on every build — these numbers are never edited by hand. Metric Value Registered companies 234… See the full description on the dataset page: https://huggingface.co/datasets/swadhinbiswas/bangladeshi-jobs.texttext-retrieval1K<n<10K0 likes272 downloads6d agoHugging Face15biplob998 /bangla-newstabular1M<n<10M0 likes258 downloads2mo agoHugging Face16kamruzzaman-asif /bangla-instruction-dataset 🧠 Bangla Instruction Dataset This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models. 📚 Dataset Splits The dataset is organized into the following splits: Split Name Source Dataset Description OdiaGenAI OdiaGenAI/all_combined_bengali_252k A large-scale collection of diverse Bangla instructions and responses. chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.texttext-generation1M<n<10M1 likes210 downloads1y agoHugging Face17FaiyazAbdullah114708 /BanglaVerse Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision–language… See the full description on the dataset page: https://huggingface.co/datasets/FaiyazAbdullah114708/BanglaVerse.imagetranslation10K<n<100K2 likes198 downloads6mo agoHugging Face18BanglaLLM /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/BanglaLLM/BanglaSafe.tabulartext-generationn<1K0 likes188 downloads1mo agoHugging Face19tareq052 /bangla-voice-03042audion<1K0 likes187 downloads2mo agoHugging Face20momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads24d agoHugging Face21Shozol /translated_gsm8k_to_bangla_traintext1K<n<10K0 likes170 downloads2y agoHugging Face22csebuetnlp /BanglaContextualBias Dataset Card for Bangla Contextual Bias The Bangla Contextual Bias dataset corresponds to the data described in the paper "An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla" accepted in ACL 2024 (Findings). Dataset Description The dataset has different parts for different bias detection experiments conducted for Bengali. WEAT & SEAT For the WEAT experiment, the dataset is translated from its English counterpart and… See the full description on the dataset page: https://huggingface.co/datasets/csebuetnlp/BanglaContextualBias.textsentence-similarityn<1K1 likes168 downloads2y agoHugging Face23Reza2kn /bangla-ocr-double-benchmark Bangla OCR Double Benchmark Two equally weighted, deterministic full-page Bangla handwriting robustness splits: bongabdo: 6,669 readability-preserving renderings balanced over all 111 Bongabdo pages. bn_htrd: 6,669 renderings balanced over all 75 actual files in the writer-separated BN-HTRd test split. These are explicitly compositional/augmentation robustness rows, not 13,338 independent writers or source documents. Every row exposes its source page ID, source SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/bangla-ocr-double-benchmark.imageimage-to-text10K<n<100K0 likes161 downloads2mo agoHugging Face24DhimanBose /Bangla_MLM_Texts_Datasettext10M<n<100M1 likes157 downloads3y agoHugging Face25md-nishat-008 /Bangla-TextBook Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu --- If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-TextBook.texttext-generation10K<n<100K2 likes148 downloads1y agoHugging Face26csebuetnlp /BanglaNMTThis is the largest Machine Translation (MT) dataset for Bengali-English, introduced in the paper `Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation`.texttranslation1M<n<10M13 likes145 downloads4y agoHugging Face27nymtheescobar /BanglaSafe BanglaSafe dataset card Overview BanglaSafe is a Bengali safety benchmark of 879 prompts covering 17 harm categories, written natively rather than translated from English. Every category is anchored to a Bangladesh statute or a documented case, and every harm instance is written five ways so that only the language and the register change. That last part is the point. Bengali is diglossic: newspaper prose and a casual text message… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/BanglaSafe.tabulartext-generationn<1K0 likes142 downloads1mo agoHugging Face28munzurul /bangla-corpus TituLM Bangla Corpus This dataset is associated with the paper TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking TituLM Bangla Corpus is one of the largest Bangla clean corpus prepared for pretraining, continual pretraining or fine-tuning Large Language Model(LLM) for improving Bangla text generation capability. This dataset contains diverse sources and categories of Bangla text. The largest part of this dataset contains filtered common crawled datasets. As we saw… See the full description on the dataset page: https://huggingface.co/datasets/munzurul/bangla-corpus.texttext-generation10M<n<100M0 likes141 downloads7mo agoHugging Face29zabir-nabil /bangla_newspaper_dataset Bangla Newspaper Dataset 400k+ bangla news samples, 25+ categories Source Data collected from https://www.prothomalo.com/archive [Copyright owned by the actual source] Github Github repository (Bi-LSTM Baseline): https://github.com/zabir-nabil/bangla-news-rnn Kaggle Version Kaggle Dataset: https://www.kaggle.com/datasets/furcifer/bangla-newspaper-dataset Inspiration The dataset can be used for Bangla text classification and generation… See the full description on the dataset page: https://huggingface.co/datasets/zabir-nabil/bangla_newspaper_dataset.tabulartext-classification100K<n<1M3 likes130 downloads2y agoHugging Face30samikhan121 /bangla_tts_iitmaudio10K<n<100K0 likes125 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.