datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
banglish_bench
BanglishBench
A smoke test for Banglish models. It answers one question: did this build break?
700 prompts, 7 categories, a floor per category, and an exit code. The same job
pytest does before you demo a feature.
It does not rank models and it does not measure quality. It tells you whether a
build is worth the time it takes to read its answers. Whether the answers are
any good still takes a person who reads Banglish.
Run it
pip install huggingface_hub
hf… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/banglish_bench.bn_en_banglish_100k_finetune
bn_en_banglish_100k_finetune
A curated, category-balanced, deduplicated 100,000-row subset of
tensorlabco/bn_en_banglish_v2
(1,370,553 source rows), built for Bangla↔English translation fine-tuning. This is the full,
3-text-field variant; see the companion
tensorlabco/bn_en_100k_finetune
for a bn_text/en_text-only projection of the exact same 100,000 rows and split.
Columns
Column
Description
category
Fine-grained topic (e.g. sports, bangladesh, book)… See the full description on the dataset page: https://huggingface.co/datasets/tensorlabco/bn_en_banglish_100k_finetune.bangla-english-banglish-pairs
Bangla-English-Banglish Trilingual Pairs
Overview
This dataset provides contrastive training pairs for fine-tuning trilingual (Bangla / Banglish / English) sentence embedding models (such as BGE-M3). It is designed to impart robustness to Banglish spelling variation.
The dataset is combined from two main sources:
LLM-generated Banglish spelling variants.
The OPUS-100 EN-BN parallel corpus.
Included Files
File
Rows
Size
Description… See the full description on the dataset page: https://huggingface.co/datasets/istiaqfuad/bangla-english-banglish-pairs.BanglishDepYT
BanglishDepYT
Banglish code-mixed YouTube comment dataset for depression-related NLP research.
Statistics
191k unlabeled comments
1k manually labeled comments
Tasks
Depression detection
Sentiment analysis
Code-mixed NLP
Data Collection stuffs
GitHub: https://github.com/Ay-on-Roy/BanglishDepYT
License
license: apache-2.0
banglish_dataset_v1banglish-speech-corpus-v0banglish-speech-corpus-v1smolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Arban221B/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: b8fa685c)
Records: 10000
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Arban221B.
Generated via Smolify.ai.
banglish_dataset_v3banking14-intents-en-bn-banglish
Multilingual Banking Intent Dataset
Dataset Overview
This dataset is a custom-built multilingual intent classification dataset designed for banking chatbot systems. It supports English, Bangla (Bengali script), and Banglish (Romanized Bengali), with limited code-mixed examples.
The dataset was created for training production-grade multilingual banking intent classifiers with strong out-of-domain fallback detection.
Dataset Size
Total Samples: 134,412 original… See the full description on the dataset page: https://huggingface.co/datasets/learn-abc/banking14-intents-en-bn-banglish.banglish-sentiment-2026Banglish Sentiment Dataset 2026 – Code-Mixed Bangla (English Script) for NLP
Synthetic dataset (~25,000 unique rows) of Banglish messages (Bangla in English letters, e.g., "Ajke onek valo lagse") labeled as positive, negative, or neutral.
Research MotivationBanglish is common in Bangladesh/South Asia for texting/social media, but few sentiment datasets exist for it. This fills the gap for code-mixed NLP, chatbots, social sentiment tools, and low-resource research.
Columns
text: Banglish… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/banglish-sentiment-2026.Banglish-Englishsmolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model AitijhyaR/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: dac3e97c)
Records: 9900
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by AitijhyaR.
Generated via Smolify.ai.
smolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Ayan-12/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 849ef9b5)
Records: 9970
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Ayan-12.
Generated via Smolify.ai.
smolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 53c8c249)
Records: 1280
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
Bangla_Banglish_hotel_booking_datasetsmolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Thorfinn05/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 1adb8318)
Records: 9960
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Thorfinn05.
Generated via Smolify.ai.
smolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Aishwarya0803/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 17808338)
Records: 9425
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Aishwarya0803.
Generated via Smolify.ai.
banglish_80K_dataset_v1polished-banglish-slmbn_en_banglish_v2smolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model ankita182005/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 475866c2)
Records: 9231
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by ankita182005.
Generated via Smolify.ai.
Banglish-Englishbanglish_dataset_pretrained_225k
