CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01findnitai /english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources. Sources: Hinglish TOP Dataset CMU English Dog HinGE PHINC source : 1 - Human Annotated , source : 0 - Synthetically Generated texttranslation100K<n<1M23 likes180 downloads3y agoHugging Face02Sujalvc /hinglish-instruct-dataset Akshar Hinglish Instruct Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns. 1. Dataset Overview Total Examples: 10,378 Base Set: 9,999 instruction-following pairs Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.texttext-generation10K<n<100K1 likes97 downloads4mo agoHugging Face03vikasaivyas /hindi-novel-sft-dataset 📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस) यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है। 🌟 प्रमुख विशेषताएँ (Key Highlights) 10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ। 100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.texttext-generationn<1K0 likes73 downloads16d agoHugging Face04AmareshHebbar /hindi-medical-sft Hindi Medical Reasoning (Medical-o1-SFT) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Medical questions → detailed chain-of-thought reasoning and clinical answers Why download this Fine-tune models for Hindi-language medical Q&A, build ABDM-compatible clinical assistants, or create multilingual medical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/hindi-medical-sft.texttext-generation10K<n<100K0 likes54 downloads3mo agoHugging Face05Nebulixlabs /Saraswati-Hindi Saraswati-Hindi Saraswati-Hindi is an English-to-Hindi parallel text dataset containing automatically translated English sentences and their corresponding Hindi translations. The dataset was created using the MyMemory Translation API to translate English text into Hindi (en → hi). It is intended for research, experimentation, and development of English-to-Hindi natural language processing (NLP) and machine translation systems. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Nebulixlabs/Saraswati-Hindi.texttext-generationn<1K0 likes48 downloads29d agoHugging Face06me-nabi /hindikrishi-farmer-advisory-dataset 🌾 HindiKrishi — Farmer Advisory Dataset 21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines. Dataset Details Detail Value Total Examples 21,069 Languages Hindi (primary), English Format JSONL (instruction, input, output) Domain Indian agriculture — crop diseases, pesticides, fertilizers, schemes License Apache 2.0 Format Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.texttext-generation10K<n<100K0 likes45 downloads2mo agoHugging Face07Ghanashyaam /CallAgentAI-Hinglish-Customer-Service CallAgent AI: Hinglish Business Conversations Dataset This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs. Why this dataset exists Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.tabulartext-generationn<1K0 likes37 downloads26d agoHugging Face08BhabhaAI /Cross-Hindi-Hinglish-chat Cross Hindi Hinglish Chat This dataset is a subset of OpenHermes where some part is converted to either Hindi or Hinglish.Note: This is in raw form. You must add "Reply in Hindi", "Reply in English" kind texts where appropriate.row_ids correspond to row id starting from 0 for OpenHermes English dataset. texttext-generation10K<n<100K1 likes35 downloads3y agoHugging Face09prakharb01 /Synthetic-Hinglish-Finetuning-Dataset Hinglish Conversations Dataset Overview This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging. Dataset Details Language: Hinglish (Hindi + English) Domain: College life, daily interactions, cultural events, and general discussions Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.texttext-generation1K<n<10K0 likes35 downloads1y agoHugging Face10Huzayfah-Patel /mindbridge-phq9-hindi-dialogues MindBridge Hindi PHQ-9/GAD-7 — Training Dialogues (2,883 rows) Single-turn ShareGPT-format dialogues for Unsloth QLoRA fine-tuning of Gemma 4 E2B. Each row: [system, user, assistant.tool_calls] where the assistant emits interpret_response({score: int 0-3, rationale_english: str, confidence: float in {0.6, 0.8, 0.95}}). Compatible with tokenizer.apply_chat_template(messages, tools=[INTERPRET_RESPONSE_TOOL_SCHEMA]) for Gemma 4 native <|tool_call> tokens. The tool schema lives in… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-dialogues.texttext-classification1K<n<10K0 likes33 downloads1mo agoHugging Face11Nitin1211 /dbpedia-hindi-cot-training-data DBpedia Hindi — Chain-of-Thought Training Data (Not Used in Final Training) 39,621 Hindi relational-triple-extraction examples in Chain-of-Thought (CoT) trace format, generated for the DBpedia Hindi Chapter (Google Summer of Code 2026), published for completeness alongside the Optimal-trace training set actually used to train the released models. Important — Not Used In The Final Model This is the exact same underlying data as the Optimal-trace training set… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-cot-training-data.texttext-generation10K<n<100K0 likes29 downloads2mo agoHugging Face12BhabhaAI /openhermes-2.5-hindi OpenHermes-2.5-Hindi ~600K rows Translated & filtered by Satpal Singh Rathore, Manav Manoj texttext-generation100K<n<1M9 likes28 downloads2y agoHugging Face13SherryT997 /HelpSteer-hinditexttext-classification1K<n<10K1 likes27 downloads3y agoHugging Face14zorqelis-ai /soreqen-hinglish SoreQen Hinglish Roman-script Hinglish conversation with an English minority slice, for training assistants that answer Indian users in the register they actually write in. Curated and published by ZorQelis AI. Train rows 36,326 Validation rows 741 Format chat messages (JSONL) Format {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}], "source": "orca_math", "lang": "english"… See the full description on the dataset page: https://huggingface.co/datasets/zorqelis-ai/soreqen-hinglish.texttext-generation10K<n<100K1 likes27 downloads1mo agoHugging Face15Nitin1211 /dbpedia-hindi-training-data DBpedia Hindi — Training Data (Relational Triple Extraction) 39,621 Hindi sentence → subject-relation-object triple examples, used to fine-tune Gemma 3 4B for the DBpedia Hindi Chapter (Google Summer of Code 2026). Format Chat-format JSONL, one example per line: { "phase": "phase1", "messages": [ {"role": "system", "content": "Extract all subject-relation-object triplets..."}, {"role": "user", "content": "<Hindi sentence>"}, {"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-training-data.texttext-generation10K<n<100K0 likes26 downloads2mo agoHugging Face16Nitin1211 /dbpedia-hindi-noisy-training-data DBpedia Hindi — Noisy Synthetic Training Data 15,581 Hindi sentence → triple examples with deliberately realistic noise, generated to support curriculum-style training for the DBpedia Hindi Chapter (Google Summer of Code 2026). Rationale Seeded from flawed (lower-scoring) examples from the original synthetic dataset, so the generated "noise" reflects genuine semantic mistakes (span boundaries, argument reversal, missing negation) rather than a weak model's… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-noisy-training-data.texttext-generation10K<n<100K0 likes24 downloads2mo agoHugging Face17thinkedgeAI /hind-promo Dataset Card: Hindi Narrative Prompt Dataset Dataset Summary This dataset consists of over 45,000 rows of Hindi language data, serving as a valuable resource for training and evaluating natural language generation models, particularly in the Hindi language domain. Each row contains the following fields: system_prompt: A detailed prompt provided in Hindi, intended to guide the generation of narratives or explanations. qas_id: Unique identifier for each question-answer… See the full description on the dataset page: https://huggingface.co/datasets/thinkedgeAI/hind-promo.texttext-generation10K<n<100K1 likes18 downloads3y agoHugging Face18bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads5mo agoHugging Face19QuantumMik /alpaca_hindi_small Alpaca Hindi Small This is a synthesized dataset created by translation of alpaca dataset from English to Hindi language. textquestion-answering1K<n<10K1 likes15 downloads3y agoHugging Face20dnyanesh /HindiMathQuestgated Overview: The HindiMathQuest: A Dataset for Mathematical Reasoning and Problem-Solving in Hindi is designed to advance the capabilities of language models in understanding and solving mathematical problems presented in the Hindi language. The dataset covers a comprehensive range of question types, including logical reasoning, numeric calculations, translation-based problems, and complex mathematical tasks typically seen in competitive exams. This dataset is intended to fill a… See the full description on the dataset page: https://huggingface.co/datasets/dnyanesh/HindiMathQuest.textquestion-answering100K<n<1M2 likes15 downloads2y agoHugging Face21marsh-mellow /hindi_wikipedia Hindi Wikipedia Corpus Dataset Description The Hindi Wikipedia Corpus is a pure Hindi text dataset derived from the Hindi-language Wikipedia (as.wikipedia.org). It contains cleaned plain text extracted from Wikipedia articles, stripped of all formatting, with non-Hindi characters completely removed. This dataset is designed for language modeling, NLP research, creating Hindi specific tokenizers, and other Hindi-language processing tasks. Data Processing… See the full description on the dataset page: https://huggingface.co/datasets/marsh-mellow/hindi_wikipedia.texttext-generation100K<n<1M0 likes15 downloads1y agoHugging Face22SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes15 downloads4mo agoHugging Face23bingbangboom /tiny-aya-translate-hinglish-casual-stripped Dataset Card for tiny-aya-translate-hinglish-casual-stripped Dataset Summary tiny-aya-translate-hinglish-casual-stripped is a lightweight, text-only derivative of the original tiny-aya-translate/hinglish-casual dataset. The original dataset is designed for simultaneous translation and contains many columns including audio references, speaker metadata, and duration. It also includes paralinguistic tags (e.g., <sigh>, <laugh>, <chuckle>) embedded within the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/tiny-aya-translate-hinglish-casual-stripped.texttext-generation10K<n<100K1 likes15 downloads3mo agoHugging Face24Nitin1211 /dbpedia-hindi-validation-data DBpedia Hindi — Validation Data (Relational Triple Extraction) 3,634 real Hindi Wikipedia sentences, held out during training, used to evaluate the fine-tuned Gemma 3 4B model for the DBpedia Hindi Chapter (Google Summer of Code 2026). Format Same chat-format JSONL as the training dataset — messages (system/user/assistant), plus score, source, trace_type fields. Composition Real Hindi Wikipedia sentences only (not synthetic), each scored ≥9/10 by an… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-validation-data.texttext-generation1K<n<10K0 likes15 downloads2mo agoHugging Face25kvrma /hindidatasettexttext-generationn<1K0 likes11 downloads2y agoHugging Face26sKT-Ai-Labs /ShORT-Hinglish_Dataset-10mgated1 🚀 ShORT-ʜɪɴɢʟɪsʜ-𝕯𝖆𝖙𝖆𝖘𝖊𝖙𝖘-𝟷𝟶ᴍ 🇮🇳 SKT AI LABS SKT AI LABS The Sovereign AI for India The Sovereign LLM Development For India (Project Surya) ✨ Overview This dataset is a monumental collection of 10 Million high-quality conversation pairs crafted in Hinglish (Hindi + English). It is meticulously engineered to… See the full description on the dataset page: https://huggingface.co/datasets/sKT-Ai-Labs/ShORT-Hinglish_Dataset-10m.texttext-generation10M<n<100M8 likes10 downloads3mo agoHugging Face27GursimranSinghBasra /nyaya-bench-hindi Nyaya Bench Hindi 500 Hindi (Devanagari) training pairs for fine-tuning LLMs on classical Indian Nyaya Panchavayava (five-limbed syllogism) reasoning. Format Each record is an instruction/output pair. All 7 reasoning fields are in Devanagari Hindi: प्रतिज्ञा (Pratijna) — Claim हेतु (Hetu) — Reason उदाहरण (Udaharana) — Example उपनय (Upanaya) — Application निगमन (Nigamana) — Conclusion पूर्वपक्ष (Purvapaksha) — Counterargument सिद्धान्त (Siddhanta) — Rebuttal… See the full description on the dataset page: https://huggingface.co/datasets/GursimranSinghBasra/nyaya-bench-hindi.texttext-generationn<1K0 likes10 downloads4mo agoHugging Face28HappyAIUser /Atma4-Hindi Dataset Card for Atma4-Hindi This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks. Dataset Description The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains: An instruction that specifies the task An optional input providing context A detailed output that addresses the instruction Usage This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma4-Hindi.texttext-generation1K<n<10K0 likes6 downloads2y agoHugging Face29Bluestrikeai /ai-hindi-chatbottexttext-generationn<1K0 likes5 downloads1y agoHugging Face30sKT-Ai-Labs /HINgated1 🚀 SKT-HIN 🇮🇳 SKT AI LABS SKT AI LABS The Sovereign AI for India The Sovereign LLM Development For India (Project Surya) ✨ Overview This dataset is a monumental collection of 320k high-quality conversation pairs crafted in Hinglish (Hindi + English). It is meticulously engineered to empower Large Language Models (LLMs) with a deep… See the full description on the dataset page: https://huggingface.co/datasets/sKT-Ai-Labs/HIN.texttext-generationn<1K3 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.