banglish
Datasets
All datasets matching “banglish”bangla-english-and-code-mixed-ecommerce-review-dataset
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
Description
The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.banglish_bench
BanglishBench
A smoke test for Banglish models. It answers one question: did this build break?
700 prompts, 7 categories, a floor per category, and an exit code. The same job
pytest does before you demo a feature.
It does not rank models and it does not measure quality. It tells you whether a
build is worth the time it takes to read its answers. Whether the answers are
any good still takes a person who reads Banglish.
Run it
pip install huggingface_hub
hf… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/banglish_bench.bn_en_banglish_100k_finetune
bn_en_banglish_100k_finetune
A curated, category-balanced, deduplicated 100,000-row subset of
tensorlabco/bn_en_banglish_v2
(1,370,553 source rows), built for Bangla↔English translation fine-tuning. This is the full,
3-text-field variant; see the companion
tensorlabco/bn_en_100k_finetune
for a bn_text/en_text-only projection of the exact same 100,000 rows and split.
Columns
Column
Description
category
Fine-grained topic (e.g. sports, bangladesh, book)… See the full description on the dataset page: https://huggingface.co/datasets/tensorlabco/bn_en_banglish_100k_finetune.bangla-english-banglish-pairs
Bangla-English-Banglish Trilingual Pairs
Overview
This dataset provides contrastive training pairs for fine-tuning trilingual (Bangla / Banglish / English) sentence embedding models (such as BGE-M3). It is designed to impart robustness to Banglish spelling variation.
The dataset is combined from two main sources:
LLM-generated Banglish spelling variants.
The OPUS-100 EN-BN parallel corpus.
Included Files
File
Rows
Size
Description… See the full description on the dataset page: https://huggingface.co/datasets/istiaqfuad/bangla-english-banglish-pairs.BanglishDepYT
BanglishDepYT
Banglish code-mixed YouTube comment dataset for depression-related NLP research.
Statistics
191k unlabeled comments
1k manually labeled comments
Tasks
Depression detection
Sentiment analysis
Code-mixed NLP
Data Collection stuffs
GitHub: https://github.com/Ay-on-Roy/BanglishDepYT
License
license: apache-2.0
Bengali-to-Banglish-Dataset
Dataset Card for Massive Bengali-English-Banglish Dictionary
Dataset Description
This is a comprehensive, massive-scale dictionary dataset containing 206,926 unique Bengali words, their corresponding English meanings, and up to 15 conversational Banglish (Latin script) transliteration variants per word.
Intended Uses & Out of Scope Use
Intended Use Cases
Machine Translation: Training neural machine translation (NMT) models to… See the full description on the dataset page: https://huggingface.co/datasets/ShayonSarker/Bengali-to-Banglish-Dataset.
