CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sifat-febo /banglish_bench BanglishBench A smoke test for Banglish models. It answers one question: did this build break? 700 prompts, 7 categories, a floor per category, and an exit code. The same job pytest does before you demo a feature. It does not rank models and it does not measure quality. It tells you whether a build is worth the time it takes to read its answers. Whether the answers are any good still takes a person who reads Banglish. &nbsp; Run it pip install huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/banglish_bench.texttext-generationn<1K0 likes152 downloads15d agoHugging Face02tensorlabco /bn_en_banglish_100k_finetune bn_en_banglish_100k_finetune A curated, category-balanced, deduplicated 100,000-row subset of tensorlabco/bn_en_banglish_v2 (1,370,553 source rows), built for Bangla↔English translation fine-tuning. This is the full, 3-text-field variant; see the companion tensorlabco/bn_en_100k_finetune for a bn_text/en_text-only projection of the exact same 100,000 rows and split. Columns Column Description category Fine-grained topic (e.g. sports, bangladesh, book)… See the full description on the dataset page: https://huggingface.co/datasets/tensorlabco/bn_en_banglish_100k_finetune.texttranslation100K<n<1M0 likes56 downloads1mo agoHugging Face03istiaqfuad /bangla-english-banglish-pairs Bangla-English-Banglish Trilingual Pairs Overview This dataset provides contrastive training pairs for fine-tuning trilingual (Bangla / Banglish / English) sentence embedding models (such as BGE-M3). It is designed to impart robustness to Banglish spelling variation. The dataset is combined from two main sources: LLM-generated Banglish spelling variants. The OPUS-100 EN-BN parallel corpus. Included Files File Rows Size Description… See the full description on the dataset page: https://huggingface.co/datasets/istiaqfuad/bangla-english-banglish-pairs.textfeature-extraction1M<n<10M0 likes39 downloads4mo agoHugging Face04AyonRoy29 /BanglishDepYT BanglishDepYT Banglish code-mixed YouTube comment dataset for depression-related NLP research. Statistics 191k unlabeled comments 1k manually labeled comments Tasks Depression detection Sentiment analysis Code-mixed NLP Data Collection stuffs GitHub: https://github.com/Ay-on-Roy/BanglishDepYT License license: apache-2.0 textn<1K0 likes34 downloads4mo agoHugging Face05kawsarahmd /banglish_dataset_v1text10K<n<100K0 likes32 downloads2y agoHugging Face06kraider24k /banglish-speech-corpus-v0audion<1K0 likes29 downloads3d agoHugging Face07kraider24k /banglish-speech-corpus-v1audion<1K0 likes29 downloads3d agoHugging Face08Arban221B /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Arban221B/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: b8fa685c) Records: 10000 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Arban221B. Generated via Smolify.ai. texttext-generation10K<n<100K0 likes23 downloads6mo agoHugging Face09kawsarahmd /banglish_dataset_v3text100K<n<1M1 likes22 downloads2y agoHugging Face10learn-abc /banking14-intents-en-bn-banglishgated Multilingual Banking Intent Dataset Dataset Overview This dataset is a custom-built multilingual intent classification dataset designed for banking chatbot systems. It supports English, Bangla (Bengali script), and Banglish (Romanized Bengali), with limited code-mixed examples. The dataset was created for training production-grade multilingual banking intent classifiers with strong out-of-domain fallback detection. Dataset Size Total Samples: 134,412 original… See the full description on the dataset page: https://huggingface.co/datasets/learn-abc/banking14-intents-en-bn-banglish.texttext-classification100K<n<1M1 likes19 downloads7mo agoHugging Face11mdsajjadullah /banglish-sentiment-2026Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral. Research MotivationBanglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research. Columns text: Banglish… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/banglish-sentiment-2026.texttext-classification10K<n<100K1 likes16 downloads7mo agoHugging Face12Ayon128 /Banglish-Englishtext10K<n<100K3 likes11 downloads2y agoHugging Face13AitijhyaR /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model AitijhyaR/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: dac3e97c) Records: 9900 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by AitijhyaR. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes11 downloads6mo agoHugging Face14Ayan-12 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Ayan-12/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 849ef9b5) Records: 9970 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Ayan-12. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes8 downloads6mo agoHugging Face15smolify /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 53c8c249) Records: 1280 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes7 downloads6mo agoHugging Face16devsajid007 /Bangla_Banglish_hotel_booking_datasettext1K<n<10K0 likes7 downloads4mo agoHugging Face17Thorfinn05 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Thorfinn05/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 1adb8318) Records: 9960 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Thorfinn05. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes5 downloads6mo agoHugging Face18Aishwarya0803 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Aishwarya0803/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 17808338) Records: 9425 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Aishwarya0803. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes5 downloads6mo agoHugging Face19kawsarahmd /banglish_80K_dataset_v1gatedtext10K<n<100K2 likes4 downloads2y agoHugging Face20ryuma007 /polished-banglish-slmtext1K<n<10K0 likes4 downloads4mo agoHugging Face21tensorlabco /bn_en_banglish_v2gatedtext1M<n<10M0 likes3 downloads1y agoHugging Face22ankita182005 /smolified-banglish-ner 🤏 smolified-banglish-ner Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model ankita182005/smolified-banglish-ner. 📦 Asset Details Origin: Smolify Foundry (Job ID: 475866c2) Records: 9231 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by ankita182005. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes3 downloads6mo agoHugging Face23Iftekhar737 /Banglish-Englishtext10K<n<100K0 likes3 downloads5mo agoHugging Face24tensorlabco /banglish_dataset_pretrained_225kgatedtext100K<n<1M0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.