CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BanglishRev /bangla-english-and-code-mixed-ecommerce-review-dataset BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce Description The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.image0 likes1.3k downloads2y agoHugging Face02sifat-febo /banglish_bench BanglishBench A smoke test for Banglish models. It answers one question: did this build break? 700 prompts, 7 categories, a floor per category, and an exit code. The same job pytest does before you demo a feature. It does not rank models and it does not measure quality. It tells you whether a build is worth the time it takes to read its answers. Whether the answers are any good still takes a person who reads Banglish. &nbsp; Run it pip install huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/banglish_bench.texttext-generationn<1K0 likes152 downloads15d agoHugging Face03tensorlabco /bn_en_banglish_100k_finetune bn_en_banglish_100k_finetune A curated, category-balanced, deduplicated 100,000-row subset of tensorlabco/bn_en_banglish_v2 (1,370,553 source rows), built for Bangla↔English translation fine-tuning. This is the full, 3-text-field variant; see the companion tensorlabco/bn_en_100k_finetune for a bn_text/en_text-only projection of the exact same 100,000 rows and split. Columns Column Description category Fine-grained topic (e.g. sports, bangladesh, book)… See the full description on the dataset page: https://huggingface.co/datasets/tensorlabco/bn_en_banglish_100k_finetune.texttranslation100K<n<1M0 likes56 downloads1mo agoHugging Face04istiaqfuad /bangla-english-banglish-pairs Bangla-English-Banglish Trilingual Pairs Overview This dataset provides contrastive training pairs for fine-tuning trilingual (Bangla / Banglish / English) sentence embedding models (such as BGE-M3). It is designed to impart robustness to Banglish spelling variation. The dataset is combined from two main sources: LLM-generated Banglish spelling variants. The OPUS-100 EN-BN parallel corpus. Included Files File Rows Size Description… See the full description on the dataset page: https://huggingface.co/datasets/istiaqfuad/bangla-english-banglish-pairs.textfeature-extraction1M<n<10M0 likes39 downloads4mo agoHugging Face05kraider24k /banglishbench-evalaudion<1K0 likes37 downloads3d agoHugging Face06AyonRoy29 /BanglishDepYT BanglishDepYT Banglish code-mixed YouTube comment dataset for depression-related NLP research. Statistics 191k unlabeled comments 1k manually labeled comments Tasks Depression detection Sentiment analysis Code-mixed NLP Data Collection stuffs GitHub: https://github.com/Ay-on-Roy/BanglishDepYT License license: apache-2.0 textn<1K0 likes34 downloads4mo agoHugging Face07kawsarahmd /banglish_dataset_v1text10K<n<100K0 likes32 downloads2y agoHugging Face08ShayonSarker /Bengali-to-Banglish-Dataset Dataset Card for Massive Bengali-English-Banglish Dictionary Dataset Description This is a comprehensive, massive-scale dictionary dataset containing 206,926 unique Bengali words, their corresponding English meanings, and up to 15 conversational Banglish (Latin script) transliteration variants per word. Intended Uses & Out of Scope Use Intended Use Cases Machine Translation: Training neural machine translation (NMT) models to… See the full description on the dataset page: https://huggingface.co/datasets/ShayonSarker/Bengali-to-Banglish-Dataset.translation100K<n<1M1 likes32 downloads2mo agoHugging Face09kraider24k /banglish-speech-corpus-v0audion<1K0 likes29 downloads3d agoHugging Face10kraider24k /banglish-speech-corpus-v1audion<1K0 likes29 downloads3d agoHugging Face11Arban221B /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Arban221B/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: b8fa685c) Records: 10000 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Arban221B. Generated via Smolify.ai. texttext-generation10K<n<100K0 likes23 downloads6mo agoHugging Face12kawsarahmd /banglish_dataset_v3text100K<n<1M1 likes22 downloads2y agoHugging Face13learn-abc /banking14-intents-en-bn-banglishgated Multilingual Banking Intent Dataset Dataset Overview This dataset is a custom-built multilingual intent classification dataset designed for banking chatbot systems. It supports English, Bangla (Bengali script), and Banglish (Romanized Bengali), with limited code-mixed examples. The dataset was created for training production-grade multilingual banking intent classifiers with strong out-of-domain fallback detection. Dataset Size Total Samples: 134,412 original… See the full description on the dataset page: https://huggingface.co/datasets/learn-abc/banking14-intents-en-bn-banglish.texttext-classification100K<n<1M1 likes19 downloads7mo agoHugging Face14mdsajjadullah /banglish-sentiment-2026Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral. Research MotivationBanglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research. Columns text: Banglish… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/banglish-sentiment-2026.texttext-classification10K<n<100K1 likes16 downloads7mo agoHugging Face15Ayon128 /Banglish-Englishtext10K<n<100K3 likes11 downloads2y agoHugging Face16AitijhyaR /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model AitijhyaR/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: dac3e97c) Records: 9900 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by AitijhyaR. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes11 downloads6mo agoHugging Face17Ayan-12 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Ayan-12/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 849ef9b5) Records: 9970 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Ayan-12. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes8 downloads6mo agoHugging Face18smolify /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 53c8c249) Records: 1280 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes7 downloads6mo agoHugging Face19devsajid007 /Bangla_Banglish_hotel_booking_datasettext1K<n<10K0 likes7 downloads4mo agoHugging Face20Thorfinn05 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Thorfinn05/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 1adb8318) Records: 9960 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Thorfinn05. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes5 downloads6mo agoHugging Face21Aishwarya0803 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Aishwarya0803/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 17808338) Records: 9425 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Aishwarya0803. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes5 downloads6mo agoHugging Face22kawsarahmd /banglish_80K_dataset_v1gatedtext10K<n<100K2 likes4 downloads2y agoHugging Face23ryuma007 /polished-banglish-slmtext1K<n<10K0 likes4 downloads4mo agoHugging Face24tensorlabco /bn_en_banglish_v2gatedtext1M<n<10M0 likes3 downloads1y agoHugging Face25ankita182005 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model ankita182005/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 475866c2) Records: 9231 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by ankita182005. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes3 downloads6mo agoHugging Face26Iftekhar737 /Banglish-Englishtext10K<n<100K0 likes3 downloads5mo agoHugging Face27tensorlabco /banglish_dataset_pretrained_225kgatedtext100K<n<1M0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.