datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources.
Sources:
Hinglish TOP Dataset
CMU English Dog
HinGE
PHINC
source : 1 - Human Annotated ,
source : 0 - Synthetically Generated
hinglish-instruct-dataset
Akshar Hinglish Instruct
Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns.
1. Dataset Overview
Total Examples: 10,378
Base Set: 9,999 instruction-following pairs
Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.hindi-novel-sft-dataset
📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस)
यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है।
🌟 प्रमुख विशेषताएँ (Key Highlights)
10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ।
100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.hindi-medical-sft
Hindi Medical Reasoning (Medical-o1-SFT)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Medical questions → detailed chain-of-thought reasoning and clinical answers
Why download this
Fine-tune models for Hindi-language medical Q&A, build ABDM-compatible clinical assistants, or create multilingual medical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/hindi-medical-sft.Saraswati-Hindi
Saraswati-Hindi
Saraswati-Hindi is an English-to-Hindi parallel text dataset containing automatically translated English sentences and their corresponding Hindi translations.
The dataset was created using the MyMemory Translation API to translate English text into Hindi (en → hi). It is intended for research, experimentation, and development of English-to-Hindi natural language processing (NLP) and machine translation systems.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Nebulixlabs/Saraswati-Hindi.hindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.CallAgentAI-Hinglish-Customer-Service
CallAgent AI: Hinglish Business Conversations Dataset
This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs.
Why this dataset exists
Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.Cross-Hindi-Hinglish-chat
Cross Hindi Hinglish Chat
This dataset is a subset of OpenHermes where some part is converted to either Hindi or Hinglish.Note: This is in raw form. You must add "Reply in Hindi", "Reply in English" kind texts where appropriate.row_ids correspond to row id starting from 0 for OpenHermes English dataset.
Synthetic-Hinglish-Finetuning-Dataset
Hinglish Conversations Dataset
Overview
This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging.
Dataset Details
Language: Hinglish (Hindi + English)
Domain: College life, daily interactions, cultural events, and general discussions
Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.mindbridge-phq9-hindi-dialogues
MindBridge Hindi PHQ-9/GAD-7 — Training Dialogues (2,883 rows)
Single-turn ShareGPT-format dialogues for Unsloth QLoRA fine-tuning of
Gemma 4 E2B. Each row: [system, user, assistant.tool_calls] where the
assistant emits interpret_response({score: int 0-3, rationale_english: str, confidence: float in {0.6, 0.8, 0.95}}). Compatible with
tokenizer.apply_chat_template(messages, tools=[INTERPRET_RESPONSE_TOOL_SCHEMA])
for Gemma 4 native <|tool_call> tokens.
The tool schema lives in… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-dialogues.dbpedia-hindi-cot-training-data
DBpedia Hindi — Chain-of-Thought Training Data (Not Used in Final Training)
39,621 Hindi relational-triple-extraction examples in Chain-of-Thought (CoT) trace format, generated for the DBpedia Hindi Chapter (Google Summer of Code 2026), published for completeness alongside the Optimal-trace training set actually used to train the released models.
Important — Not Used In The Final Model
This is the exact same underlying data as the Optimal-trace training set… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-cot-training-data.openhermes-2.5-hindi
OpenHermes-2.5-Hindi
~600K rows Translated & filtered by Satpal Singh Rathore, Manav Manoj
HelpSteer-hindisoreqen-hinglish
SoreQen Hinglish
Roman-script Hinglish conversation with an English minority slice, for training assistants that answer Indian users in the register they actually write in.
Curated and published by ZorQelis AI.
Train rows
36,326
Validation rows
741
Format
chat messages (JSONL)
Format
{"messages": [{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}],
"source": "orca_math", "lang": "english"… See the full description on the dataset page: https://huggingface.co/datasets/zorqelis-ai/soreqen-hinglish.dbpedia-hindi-training-data
DBpedia Hindi — Training Data (Relational Triple Extraction)
39,621 Hindi sentence → subject-relation-object triple examples, used to fine-tune Gemma 3 4B for the DBpedia Hindi Chapter (Google Summer of Code 2026).
Format
Chat-format JSONL, one example per line:
{
"phase": "phase1",
"messages": [
{"role": "system", "content": "Extract all subject-relation-object triplets..."},
{"role": "user", "content": "<Hindi sentence>"},
{"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-training-data.dbpedia-hindi-noisy-training-data
DBpedia Hindi — Noisy Synthetic Training Data
15,581 Hindi sentence → triple examples with deliberately realistic noise, generated to support curriculum-style training for the DBpedia Hindi Chapter (Google Summer of Code 2026).
Rationale
Seeded from flawed (lower-scoring) examples from the original synthetic dataset, so the generated "noise" reflects genuine semantic mistakes (span boundaries, argument reversal, missing negation) rather than a weak model's… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-noisy-training-data.hind-promo
Dataset Card: Hindi Narrative Prompt Dataset
Dataset Summary
This dataset consists of over 45,000 rows of Hindi language data, serving as a valuable resource for training and evaluating natural language generation models, particularly in the Hindi language domain. Each row contains the following fields:
system_prompt: A detailed prompt provided in Hindi, intended to guide the generation of narratives or explanations.
qas_id: Unique identifier for each question-answer… See the full description on the dataset page: https://huggingface.co/datasets/thinkedgeAI/hind-promo.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.alpaca_hindi_small
Alpaca Hindi Small
This is a synthesized dataset created by translation of alpaca dataset from English to Hindi language.
HindiMathQuest
Overview:
The HindiMathQuest: A Dataset for Mathematical Reasoning and Problem-Solving in Hindi is designed to advance the capabilities of language models in understanding and solving mathematical problems presented in the Hindi language. The dataset covers a comprehensive range of question types, including logical reasoning, numeric calculations, translation-based problems, and complex mathematical tasks typically seen in competitive exams. This dataset is intended to fill a… See the full description on the dataset page: https://huggingface.co/datasets/dnyanesh/HindiMathQuest.hindi_wikipedia
Hindi Wikipedia Corpus
Dataset Description
The Hindi Wikipedia Corpus is a pure Hindi text dataset derived from the Hindi-language Wikipedia (as.wikipedia.org).
It contains cleaned plain text extracted from Wikipedia articles, stripped of all formatting, with non-Hindi characters completely removed.
This dataset is designed for language modeling, NLP research, creating Hindi specific tokenizers, and other Hindi-language processing tasks.
Data Processing… See the full description on the dataset page: https://huggingface.co/datasets/marsh-mellow/hindi_wikipedia.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.tiny-aya-translate-hinglish-casual-stripped
Dataset Card for tiny-aya-translate-hinglish-casual-stripped
Dataset Summary
tiny-aya-translate-hinglish-casual-stripped is a lightweight, text-only derivative of the original tiny-aya-translate/hinglish-casual dataset.
The original dataset is designed for simultaneous translation and contains many columns including audio references, speaker metadata, and duration. It also includes paralinguistic tags (e.g., <sigh>, <laugh>, <chuckle>) embedded within the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/tiny-aya-translate-hinglish-casual-stripped.dbpedia-hindi-validation-data
DBpedia Hindi — Validation Data (Relational Triple Extraction)
3,634 real Hindi Wikipedia sentences, held out during training, used to evaluate the fine-tuned Gemma 3 4B model for the DBpedia Hindi Chapter (Google Summer of Code 2026).
Format
Same chat-format JSONL as the training dataset — messages (system/user/assistant), plus score, source, trace_type fields.
Composition
Real Hindi Wikipedia sentences only (not synthetic), each scored ≥9/10 by an… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-validation-data.hindidatasetShORT-Hinglish_Dataset-10m1
🚀 ShORT-ʜɪɴɢʟɪsʜ-𝕯𝖆𝖙𝖆𝖘𝖊𝖙𝖘-𝟷𝟶ᴍ 🇮🇳
SKT AI LABS
SKT AI LABS
The Sovereign AI for India
The Sovereign LLM Development For India (Project Surya)
✨ Overview
This dataset is a monumental collection of 10 Million high-quality conversation pairs crafted in Hinglish (Hindi + English). It is meticulously engineered to… See the full description on the dataset page: https://huggingface.co/datasets/sKT-Ai-Labs/ShORT-Hinglish_Dataset-10m.nyaya-bench-hindi
Nyaya Bench Hindi
500 Hindi (Devanagari) training pairs for fine-tuning LLMs on classical Indian Nyaya Panchavayava (five-limbed syllogism) reasoning.
Format
Each record is an instruction/output pair. All 7 reasoning fields are in Devanagari Hindi:
प्रतिज्ञा (Pratijna) — Claim
हेतु (Hetu) — Reason
उदाहरण (Udaharana) — Example
उपनय (Upanaya) — Application
निगमन (Nigamana) — Conclusion
पूर्वपक्ष (Purvapaksha) — Counterargument
सिद्धान्त (Siddhanta) — Rebuttal… See the full description on the dataset page: https://huggingface.co/datasets/GursimranSinghBasra/nyaya-bench-hindi.Atma4-Hindi
Dataset Card for Atma4-Hindi
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma4-Hindi.ai-hindi-chatbotHIN1
🚀 SKT-HIN 🇮🇳
SKT AI LABS
SKT AI LABS
The Sovereign AI for India
The Sovereign LLM Development For India (Project Surya)
✨ Overview
This dataset is a monumental collection of 320k high-quality conversation pairs crafted in Hinglish (Hindi + English). It is meticulously engineered to empower Large Language Models (LLMs) with a deep… See the full description on the dataset page: https://huggingface.co/datasets/sKT-Ai-Labs/HIN.
