CoolFace
Datasetpublic

Yugrathee28/Hinglish-dataset

๐Ÿ‡ฎ๐Ÿ‡ณ Hinglish Dataset โ€” 1.4 Million Samples Industrial-Grade Code-Mixed NLP Dataset | By ScaleIndia AI ยท Founder: Yug Rathee This repository contains a 5,000-row teaser sample from the full 1.46 Million+ Hinglish comment dataset built by Scaling YUG (Founder: Yug Rathee(yugrathee28@gmail.com)). Provided strictly for research and evaluation purposes only. Commercial use, redistribution, or production-model training requires explicit written consent fromโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Yugrathee28/Hinglish-dataset.

sourceHugging Faceupdated 5mo agoView on Hugging Face
2likes59downloads
Dataset Card

๐Ÿ‡ฎ๐Ÿ‡ณ Hinglish Dataset โ€” 1.4 Million Samples

Industrial-Grade Code-Mixed NLP Dataset | By ScaleIndia AI ยท Founder: Yug Rathee

Language-orange) Rows Teaser Pipeline Quality License Contact


This repository contains a 5,000-row teaser sample from the full 1.46 Million+ Hinglish comment dataset built by Scaling YUG (Founder: Yug Rathee(yugrathee28@gmail.com)). Provided strictly for research and evaluation purposes only. Commercial use, redistribution, or production-model training requires explicit written consent from ScaleIndia AI.

๐Ÿ“Œ What is Hinglish?

Hinglish is a naturally spoken blend of Hindi and English โ€” the dominant code-mixed language used by 500 Million+ people across India's internet. It appears in YouTube comments, WhatsApp chats, Twitter/X posts, and product reviews.

It is one of the most underrepresented yet commercially valuable languages for AI training today. Most LLMs and NLP models perform poorly on Hinglish because no large, clean, labeled dataset existed โ€” until now.


๐Ÿ”ท Full Dataset Overview

FieldValue
Total Rows (Full)1,466,926
Teaser Rows (This File)5,000
LanguageHinglish (Hindi + English code-mixed)
SourceYouTube comments (industrial-scale scrape)
Pipeline Versionv7.0-turbo
FormatJSON Array / JSONL (UTF-8)
Built ByScaleIndia AI โ€” Founder: Yug Rathee

โš™๏ธ Full Pipeline Cleaning Report โ€” v7.0-turbo

The full 1.46M dataset was processed through an industrial-grade cleaning pipeline. Below is the verified pipeline output report.
======================================================================
  HINGLISH TURBO CLEANING REPORT  โ€”  v7.0-turbo
  Config hash : 37f130be5a81f15b
  Fast JSON   : orjson
======================================================================
  Runtime     : 43.0 min  (2,582 s)
  Throughput  : 583 rows/s
  Dedup engine: MinHash LSH (full dataset)
======================================================================

๐Ÿ“ฆ Row Count Summary

MetricCount%
Total Input Rows1,506,178100%
Skipped (checkpoint)0โ€”
Actually Processed1,506,178100%
โœ… KEPT (clean)1,466,92697.4%
๐Ÿ—‘๏ธ REMOVED (trash)39,2522.6%
97.4% retention rate โ€” the input data was already high quality. Only genuinely noisy rows were removed.

๐Ÿ—‘๏ธ Removal Breakdown

CategoryRemoved
SCHEMA
Invalid schema rows0
NOISE / SPAM
Emoji-only spam0
PII-only rows0
Garbage characters0
Word repetition spam7
Random number spam1
Empty / blank0
Subtotal8
LANGUAGE
Pure Devanagari Hindi0
Pure English0
Mostly English (>50%)29,670
Broken translation1,589
Language spam0
Subtotal31,259
QUALITY
Abusive content5,778
Low quality score0
DUPLICATES
Exact duplicates2,207
Near duplicates (MinHash LSH)0
Subtotal2,207
TOTAL REMOVED39,252

๐Ÿท๏ธ Full Dataset Label Quality

Signal StrengthCount%
Strong signal labels230,24115.7%
Weak signal labels543,61137.1%
Fallback labels693,07447.2%
Short text rows (โ‰ค5 words, flagged)250,79117.1%

๐Ÿ“Š Quality Score Histogram (Kept Rows)

[0.0โ€“0.1]                    0
[0.1โ€“0.2]                    0
[0.2โ€“0.3]                    0
[0.3โ€“0.4]                    0
[0.4โ€“0.5]                    3
[0.5โ€“0.6]                   96
[0.6โ€“0.7]                2,854
[0.7โ€“0.8] โ–ˆ             63,377
[0.8โ€“0.9] โ–ˆโ–ˆโ–ˆ           225,940
[0.9โ€“1.0] โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ 1,174,656
80%+ of kept rows score 0.9โ€“1.0 quality โ€” the dataset is overwhelmingly high-grade text.

๐Ÿ“‹ 5,000-Row Teaser โ€” Audit Report

This teaser file (hinglish_teaser_5k.json) is a stratified, confidence-biased sample of the full dataset.

๐Ÿ“ Text Quality

MetricValue
Average Word Count16.4 words/row
Max Word Count229 words
Lexical Diversity16.10% (unique/total word ratio)

๐Ÿท๏ธ Label Reliability

MetricValue
Strong Signal Rows2,002 (40.0%)
Average Label Confidence0.61 / 1.00
Intent Entropy2.31 (High โ€” max theoretical ~3.17)
Label MethodCount%Meaning
strong_signal2,00240.0%Multiple strong keyword matches โ€” highly reliable
weak_signal1,33026.6%At least one signal matched โ€” generally reliable
fallback1,66833.4%Labeled Neutral by elimination

๐ŸŽฏ Intent Distribution (5k Teaser)

IntentCount%
Neutral1,43228.6%
Question1,15123.0%
Request1,12722.5%
Appreciation98419.7%
Criticism1122.2%
Humor961.9%
Complaint781.6%
Suggestion180.4%
Sarcasm20.04%
Total5,000100%

๐Ÿ˜Š Emotion Distribution (5k Teaser)

EmotionCount%
Neutral2,61152.2%
Happy1,06421.3%
Curious58311.7%
Frustrated1863.7%
Sad1793.6%
Humor1402.8%
Fear901.8%
Surprised711.4%
Angry621.2%
Disgusted140.3%

โœ… Teaser Health Checks

CheckResult
No duplicate rowsโœ… Pass
IDs sequential 1โ€“5000โœ… Pass
All 9 intent classes presentโœ… Pass
All 10 emotion classes presentโœ… Pass
Intent entropy > 2.0 (high diversity)โœ… Pass โ€” 2.31
Avg confidence > 0.5โœ… Pass โ€” 0.61
First rows at peak confidenceโœ… Pass โ€” 1.00
PII scrubbedโœ… Pass
Abusive content filteredโœ… Pass

๐Ÿ“ Data Schema

Every row follows this exact structure:

json
{
    "id": 1,
    "text": "Yaar ye video bahut zyada helpful thi, seriously thank you so much!",
    "intent": "Appreciation",
    "emotion": "Happy",
    "toxicity": "Low",
    "sarcasm": "No",
    "language": "hinglish",
    "quality_score": 0.82,
    "label_confidence": 0.9,
    "label_method": "strong_signal",
    "is_short": false
}
FieldTypeDescription
idIntegerUnique sequential row ID
textStringCleaned, PII-scrubbed Hinglish text
intentStringOne of 9 intent classes
emotionStringOne of 10 emotion classes
toxicityStringLow / Medium / High
sarcasmStringYes / No
languageStringAlways "hinglish"
quality_scoreFloatHeuristic text quality score (0.0โ€“1.0)
label_confidenceFloatLabeling confidence score (0.0โ€“1.0)
label_methodStringstrong_signal / weak_signal / fallback
is_shortBooleantrue if fewer than 6 training-grade words

๐Ÿ”ฌ How This Dataset Was Built

  1. 1.Scraping โ€” YouTube comments collected via industrial async scraper at scale
  2. 2.Normalization โ€” Unicode normalization, emoji handling, encoding fixes
  3. 3.PII Scrubbing โ€” Phone numbers, emails, handle names removed
  4. 4.Noise Filtering โ€” Emoji-only, word repetition spam, garbage chars removed
  5. 5.Language Filtering โ€” Pure Hindi, pure English, broken translations excluded
  6. 6.Quality Scoring โ€” Heuristic scoring on lexical richness, structure, length
  7. 7.Deduplication โ€” Exact SHA-256 + near-duplicate MinHash LSH (128 permutations)
  8. 8.Labeling โ€” Mega-Regex heuristic labeler across 9 intent + 10 emotion classes
  9. 9.Confidence Scoring โ€” Every row assigned label_confidence (0.0โ€“1.0)

๐Ÿ’ผ Ideal Use Cases

  • โ€”๐Ÿค– Conversational AI & Chatbot fine-tuning โ€” intent + emotion labels ready to use
  • โ€”๐Ÿง  Sentiment & emotion analysis in code-mixed Hinglish
  • โ€”๐Ÿ“š Low-resource NLP research โ€” one of the largest labeled Hinglish datasets publicly available
  • โ€”๐Ÿ” Sarcasm & toxicity classification benchmarks
  • โ€”๐Ÿ—๏ธ Pre-training or fine-tuning multilingual / code-mixed LLMs
  • โ€”๐Ÿ“Š Academic research on South Asian internet language

โš–๏ธ License & Legal

Copyright ยฉ 2026 ScaleIndia AI โ€” Yug Rathee. All Rights Reserved.

This dataset is provided STRICTLY for research and evaluation purposes only.

PERMITTED:
  โœ… Viewing and evaluating dataset quality
  โœ… Academic and non-commercial research with attribution
  โœ… Sharing with proper credit to ScaleIndia AI (Yug Rathee)

STRICTLY PROHIBITED WITHOUT EXPLICIT WRITTEN CONSENT:
  โŒ Commercial use of any kind
  โŒ Redistribution, re-uploading, or mirroring this data
  โŒ Training production-grade or commercial AI/ML models
  โŒ Selling, licensing, or sublicensing to third parties
  โŒ Claiming ownership or authorship of this dataset

Violation of these terms may result in legal action under applicable
copyright and intellectual property law.

๐Ÿ’ฐ Purchase the Full Dataset

The full 1,466,926-row Hinglish dataset is available for commercial licensing.

What you get with the full dataset:

  • โ€”โœ… 1,466,926 cleaned, labeled, deduplicated rows
  • โ€”โœ… 9 intent classes + 10 emotion classes on every row
  • โ€”โœ… Toxicity & sarcasm flags
  • โ€”โœ… Quality score + label confidence score
  • โ€”โœ… JSON or JSONL format (your choice)
  • โ€”โœ… Commercial training rights (terms apply)
  • โ€”โœ… Support for dataset integration queries

๐Ÿ“ฌ Contact ScaleIndia AI

PlatformContact
Emailyugrathee28@gmail.com
Instagram@yugrathee.xe
For enterprise licensing, custom dataset slices, or bulk purchases โ€” email with your use case and company name for a quote.

Built with โค๏ธ by Yug Rathee โ€” ScaleIndia AI โ€” April 2026