CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Abhishekcr448 /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M20 likes328 downloads2y agoHugging Face02findnitai /english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources. Sources: Hinglish TOP Dataset CMU English Dog HinGE PHINC source : 1 - Human Annotated , source : 0 - Synthetically Generated texttranslation100K<n<1M23 likes171 downloads3y agoHugging Face03rvv-karma /English-Hinglish-TOP English Hinglish (TOP Dataset) This dataset is generated from Hinglish-TOP Dataset. Data distribution: Train a. Human Generated - 6513 b. Synthetically generated - 170083 Validation a. Human Generated - 1390 b. Synthetically generated - 0 Test a. Human Generated - 6513 b. Synthetically generated - 0 texttranslation100K<n<1M1 likes131 downloads3y agoHugging Face04Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes124 downloads9mo agoHugging Face05shreyansh12183 /shreyansh-hinglish-english-stem-500k 🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations. 📖 Overview In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.textquestion-answering100K<n<1M0 likes104 downloads3d agoHugging Face06Sujalvc /hinglish-instruct-dataset Akshar Hinglish Instruct Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns. 1. Dataset Overview Total Examples: 10,378 Base Set: 9,999 instruction-following pairs Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.texttext-generation10K<n<100K1 likes99 downloads4mo agoHugging Face07rvv-karma /English-Hinglish English Hinglish English to Hinglish Dataset processed from findnitai/english-to-hinglish. Sources: Hinglish TOP Dataset CMU English Dog HinGE PHINC texttranslation100K<n<1M0 likes59 downloads3y agoHugging Face08saidutta69 /hinglish-bench Hinglish-Bench 📄 Paper: Hinglish-Bench — Reference-Free Benchmark for LLM Hinglish Text Generation (gist preprint) A reference-free benchmark for measuring how well LLMs generate natural Roman-script Hinglish — the Hindi-English code-mixing that hundreds of millions of Indians actually speak, type, and read online. Reference-free by design. Hinglish has no canonical spelling and no single "correct" rendering, so there are no gold references and no BLEU. Quality is… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/hinglish-bench.texttext-generationn<1K0 likes53 downloads1mo agoHugging Face09codebyam /Hinglish-Hindi-Transliteration-Datasetgated Hinglish-Hindi Transliteration Dataset We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.texttext-generation1K<n<10K2 likes44 downloads1y agoHugging Face10smangrul /hinglish_self_instruct_v0 Hinglish Instruct Dataset using Self Instruct method The prompt used for generating the samples: You are asked to come up with a set of 50 diverse task instructions in Hinglish or Hindi. These task instructions will be given to a GPT model and we will evaluate the GPT model for completing the instructions. Here are the requirements: 1. Try not to repeat the verb for each instruction to maximize diversity. 2. The language used for the instruction also should be diverse. For example… See the full description on the dataset page: https://huggingface.co/datasets/smangrul/hinglish_self_instruct_v0.texttext-generation1K<n<10K8 likes42 downloads3y agoHugging Face11Ghanashyaam /CallAgentAI-Hinglish-Customer-Service CallAgent AI: Hinglish Business Conversations Dataset This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs. Why this dataset exists Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.tabulartext-generationn<1K0 likes37 downloads24d agoHugging Face12BhabhaAI /Cross-Hindi-Hinglish-chat Cross Hindi Hinglish Chat This dataset is a subset of OpenHermes where some part is converted to either Hindi or Hinglish.Note: This is in raw form. You must add "Reply in Hindi", "Reply in English" kind texts where appropriate.row_ids correspond to row id starting from 0 for OpenHermes English dataset. texttext-generation10K<n<100K1 likes36 downloads3y agoHugging Face13prakharb01 /Synthetic-Hinglish-Finetuning-Dataset Hinglish Conversations Dataset Overview This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging. Dataset Details Language: Hinglish (Hindi + English) Domain: College life, daily interactions, cultural events, and general discussions Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.texttext-generation1K<n<10K0 likes34 downloads1y agoHugging Face14zorqelis-ai /soreqen-hinglish SoreQen Hinglish Roman-script Hinglish conversation with an English minority slice, for training assistants that answer Indian users in the register they actually write in. Curated and published by ZorQelis AI. Train rows 36,326 Validation rows 741 Format chat messages (JSONL) Format {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}], "source": "orca_math", "lang": "english"… See the full description on the dataset page: https://huggingface.co/datasets/zorqelis-ai/soreqen-hinglish.texttext-generation10K<n<100K1 likes32 downloads1mo agoHugging Face15hbpkillerX /alpaca-cleaned-hinglish Alpaca Cleaned (Hinglish Version) Dataset Description This is a high-quality Hinglish (Hindi written in Latin script) translation of the yahma/alpaca-cleaned dataset. It is designed for instruction fine-tuning large language models to make them conversational in Indian contexts. Dataset Summary Original Source: yahma/alpaca-cleaned (51,760 rows) Language: Hinglish (Code mixed Hindi-English) Translation Method: High-precision batch translation using… See the full description on the dataset page: https://huggingface.co/datasets/hbpkillerX/alpaca-cleaned-hinglish.texttext-generation10K<n<100K1 likes28 downloads8mo agoHugging Face16VaniAgent /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/VaniAgent/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M1 likes26 downloads7mo agoHugging Face17Subh775 /formatted-hindi-hinglish-cot Formatted Hindi-Hinglish Chain-of-Thought Dataset This is the reformatted version of the adi-kmt/hindi-hinglish-cot dataset, structured in the Alpaca instruction format for instruction tuning language models. Original Dataset The original dataset features Chain-of-Thought (CoT) conversations in Hindi-Hinglish, with: Complex user queries in Hindi-Hinglish Assistant responses that include explicit thinking steps (marked with <think> tags) Detailed explanations in… See the full description on the dataset page: https://huggingface.co/datasets/Subh775/formatted-hindi-hinglish-cot.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face18nooruiit-864 /hinglish-ai-ml-tutor-dataset Hinglish AI/ML Tutor Dataset Dataset Description A hand-curated instruction-tuning dataset of 200 Q&A pairs covering AI/ML engineering concepts (tokenization, embeddings, transformers, RAG, LoRA/QLoRA, STT/TTS, deployment, and web security). Every answer follows a consistent "Hinglish tutor" persona: an everyday analogy first, followed by the technical explanation. Format ChatML format (messages field with system/user/assistant roles), one JSON… See the full description on the dataset page: https://huggingface.co/datasets/nooruiit-864/hinglish-ai-ml-tutor-dataset.texttext-generationn<1K0 likes20 downloads1mo agoHugging Face19SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads4mo agoHugging Face20bingbangboom /tiny-aya-translate-hinglish-casual-stripped Dataset Card for tiny-aya-translate-hinglish-casual-stripped Dataset Summary tiny-aya-translate-hinglish-casual-stripped is a lightweight, text-only derivative of the original tiny-aya-translate/hinglish-casual dataset. The original dataset is designed for simultaneous translation and contains many columns including audio references, speaker metadata, and duration. It also includes paralinguistic tags (e.g., <sigh>, <laugh>, <chuckle>) embedded within the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/tiny-aya-translate-hinglish-casual-stripped.texttext-generation10K<n<100K1 likes16 downloads3mo agoHugging Face21bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes15 downloads5mo agoHugging Face22buggiebug /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/buggiebug/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M0 likes13 downloads8mo agoHugging Face23Bluestrikeai /Conversational-Hinglishtexttext-generationn<1K1 likes12 downloads1y agoHugging Face24a1b8h04i /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/a1b8h04i/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M0 likes11 downloads4mo agoHugging Face25sKT-Ai-Labs /ShORT-Hinglish_Dataset-10mgated1 🚀 ShORT-ʜɪɴɢʟɪsʜ-𝕯𝖆𝖙𝖆𝖘𝖊𝖙𝖘-𝟷𝟶ᴍ 🇮🇳 SKT AI LABS SKT AI LABS The Sovereign AI for India The Sovereign LLM Development For India (Project Surya) ✨ Overview This dataset is a monumental collection of 10 Million high-quality conversation pairs crafted in Hinglish (Hindi + English). It is meticulously engineered to… See the full description on the dataset page: https://huggingface.co/datasets/sKT-Ai-Labs/ShORT-Hinglish_Dataset-10m.texttext-generation10M<n<100M8 likes10 downloads3mo agoHugging Face26hyperneuronAILabs /Hinglishqna-llm-tunegatedThis dataset is hinglish qna ,which involves abusive language from customers in financial domain. It also involves emotional cues such as , etc. textquestion-answering10K<n<100K1 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.