datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources.
Sources:
Hinglish TOP Dataset
CMU English Dog
HinGE
PHINC
source : 1 - Human Annotated ,
source : 0 - Synthetically Generated
English-Hinglish-TOP
English Hinglish (TOP Dataset)
This dataset is generated from Hinglish-TOP Dataset.
Data distribution:
Train a. Human Generated - 6513 b. Synthetically generated - 170083
Validation a. Human Generated - 1390 b. Synthetically generated - 0
Test a. Human Generated - 6513 b. Synthetically generated - 0
Hinglish_Dataset_instruction_and_rawshreyansh-hinglish-english-stem-500k
🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus
Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations.
📖 Overview
In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.hinglish-instruct-dataset
Akshar Hinglish Instruct
Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns.
1. Dataset Overview
Total Examples: 10,378
Base Set: 9,999 instruction-following pairs
Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.English-Hinglish
English Hinglish
English to Hinglish Dataset processed from findnitai/english-to-hinglish.
Sources:
Hinglish TOP Dataset
CMU English Dog
HinGE
PHINC
hinglish-bench
Hinglish-Bench
📄 Paper: Hinglish-Bench — Reference-Free Benchmark for LLM Hinglish Text
Generation
(gist preprint)
A reference-free benchmark for measuring how well LLMs generate natural
Roman-script Hinglish — the Hindi-English code-mixing that hundreds of
millions of Indians actually speak, type, and read online.
Reference-free by design. Hinglish has no canonical spelling and no single
"correct" rendering, so there are no gold references and no BLEU. Quality is… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/hinglish-bench.Hinglish-Hindi-Transliteration-Dataset
Hinglish-Hindi Transliteration Dataset
We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.hinglish_self_instruct_v0
Hinglish Instruct Dataset using Self Instruct method
The prompt used for generating the samples:
You are asked to come up with a set of 50 diverse task instructions in Hinglish or Hindi.
These task instructions will be given to a GPT model and we will evaluate the GPT model for completing the instructions.
Here are the requirements:
1. Try not to repeat the verb for each instruction to maximize diversity.
2. The language used for the instruction also should be diverse. For example… See the full description on the dataset page: https://huggingface.co/datasets/smangrul/hinglish_self_instruct_v0.CallAgentAI-Hinglish-Customer-Service
CallAgent AI: Hinglish Business Conversations Dataset
This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs.
Why this dataset exists
Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.Cross-Hindi-Hinglish-chat
Cross Hindi Hinglish Chat
This dataset is a subset of OpenHermes where some part is converted to either Hindi or Hinglish.Note: This is in raw form. You must add "Reply in Hindi", "Reply in English" kind texts where appropriate.row_ids correspond to row id starting from 0 for OpenHermes English dataset.
Synthetic-Hinglish-Finetuning-Dataset
Hinglish Conversations Dataset
Overview
This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging.
Dataset Details
Language: Hinglish (Hindi + English)
Domain: College life, daily interactions, cultural events, and general discussions
Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.soreqen-hinglish
SoreQen Hinglish
Roman-script Hinglish conversation with an English minority slice, for training assistants that answer Indian users in the register they actually write in.
Curated and published by ZorQelis AI.
Train rows
36,326
Validation rows
741
Format
chat messages (JSONL)
Format
{"messages": [{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}],
"source": "orca_math", "lang": "english"… See the full description on the dataset page: https://huggingface.co/datasets/zorqelis-ai/soreqen-hinglish.alpaca-cleaned-hinglish
Alpaca Cleaned (Hinglish Version)
Dataset Description
This is a high-quality Hinglish (Hindi written in Latin script) translation of the yahma/alpaca-cleaned dataset.
It is designed for instruction fine-tuning large language models to make them conversational in Indian contexts.
Dataset Summary
Original Source: yahma/alpaca-cleaned (51,760 rows)
Language: Hinglish (Code mixed Hindi-English)
Translation Method: High-precision batch translation using… See the full description on the dataset page: https://huggingface.co/datasets/hbpkillerX/alpaca-cleaned-hinglish.Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/VaniAgent/Hinglish-Everyday-Conversations-1M.formatted-hindi-hinglish-cot
Formatted Hindi-Hinglish Chain-of-Thought Dataset
This is the reformatted version of the adi-kmt/hindi-hinglish-cot dataset, structured in the Alpaca instruction format for instruction tuning language models.
Original Dataset
The original dataset features Chain-of-Thought (CoT) conversations in Hindi-Hinglish, with:
Complex user queries in Hindi-Hinglish
Assistant responses that include explicit thinking steps (marked with <think> tags)
Detailed explanations in… See the full description on the dataset page: https://huggingface.co/datasets/Subh775/formatted-hindi-hinglish-cot.hinglish-ai-ml-tutor-dataset
Hinglish AI/ML Tutor Dataset
Dataset Description
A hand-curated instruction-tuning dataset of 200 Q&A pairs covering AI/ML engineering
concepts (tokenization, embeddings, transformers, RAG, LoRA/QLoRA, STT/TTS, deployment,
and web security). Every answer follows a consistent "Hinglish tutor" persona: an
everyday analogy first, followed by the technical explanation.
Format
ChatML format (messages field with system/user/assistant roles), one JSON… See the full description on the dataset page: https://huggingface.co/datasets/nooruiit-864/hinglish-ai-ml-tutor-dataset.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.tiny-aya-translate-hinglish-casual-stripped
Dataset Card for tiny-aya-translate-hinglish-casual-stripped
Dataset Summary
tiny-aya-translate-hinglish-casual-stripped is a lightweight, text-only derivative of the original tiny-aya-translate/hinglish-casual dataset.
The original dataset is designed for simultaneous translation and contains many columns including audio references, speaker metadata, and duration. It also includes paralinguistic tags (e.g., <sigh>, <laugh>, <chuckle>) embedded within the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/tiny-aya-translate-hinglish-casual-stripped.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/buggiebug/Hinglish-Everyday-Conversations-1M.Conversational-HinglishHinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/a1b8h04i/Hinglish-Everyday-Conversations-1M.ShORT-Hinglish_Dataset-10m1
🚀 ShORT-ʜɪɴɢʟɪsʜ-𝕯𝖆𝖙𝖆𝖘𝖊𝖙𝖘-𝟷𝟶ᴍ 🇮🇳
SKT AI LABS
SKT AI LABS
The Sovereign AI for India
The Sovereign LLM Development For India (Project Surya)
✨ Overview
This dataset is a monumental collection of 10 Million high-quality conversation pairs crafted in Hinglish (Hindi + English). It is meticulously engineered to… See the full description on the dataset page: https://huggingface.co/datasets/sKT-Ai-Labs/ShORT-Hinglish_Dataset-10m.Hinglishqna-llm-tuneThis dataset is hinglish qna ,which involves abusive language from customers in financial domain. It also involves emotional cues such as , etc.
