CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pfin123 /hindi-aggregatedtext100K<n<1M2 likes649 downloads4y agoHugging Face02Sheeba2026 /bharatvani-hindi-speech-corpusgated BharatVani Hindi Speech Corpus (150-Hour Studio Dataset) Proprietary Speech Asset • TheCreatorOS • BharatVani AI 1. Overview The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi. Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.audiotext-to-speech100K<n<1M1 likes285 downloads6d agoHugging Face03findnitai /english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources. Sources: Hinglish TOP Dataset CMU English Dog HinGE PHINC source : 1 - Human Annotated , source : 0 - Synthetically Generated texttranslation100K<n<1M23 likes180 downloads3y agoHugging Face04inverse-scaling /hindsight-neglect-10shot inverse-scaling/hindsight-neglect-10shot (‘The Floating Droid’) General description This task tests whether language models are able to assess whether a bet was worth taking based on its expected value. The author provides few shot examples in which the model predicts whether a bet is worthwhile by correctly answering yes or no when the expected value of the bet is positive (where the model should respond that ‘yes’, taking the bet is the right decision) or negative (‘no’… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling/hindsight-neglect-10shot.textmultiple-choicen<1K5 likes138 downloads4y agoHugging Face05QCRI /LlamaLens-Hindi LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi.texttext-classification100K<n<1M1 likes117 downloads2y agoHugging Face06FreedomIntelligence /alpaca-gpt4-hindiThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K1 likes115 downloads3y agoHugging Face07FreedomIntelligence /evol-instruct-hindiThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K2 likes111 downloads3y agoHugging Face08QCRI /LlamaLens-Hindi-Native LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi-Native.texttext-classification100K<n<1M0 likes105 downloads2y agoHugging Face09Sujalvc /hinglish-instruct-dataset Akshar Hinglish Instruct Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns. 1. Dataset Overview Total Examples: 10,378 Base Set: 9,999 instruction-following pairs Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.texttext-generation10K<n<100K1 likes97 downloads4mo agoHugging Face10nirantk /chaii-hindi-and-tamil-question-answeringtextquestion-answering1K<n<10K0 likes86 downloads3y agoHugging Face11AIhnIndicRag /Hindi_Fevertexttext-retrieval1M<n<10M0 likes86 downloads2y agoHugging Face12OdiaGenAI /sentiment_analysis_hindiConventions followed to decide the polarity: - labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'. labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.texttext-classification1K<n<10K2 likes80 downloads3y agoHugging Face13rmahesh /UP_CET_Hindi_examstextquestion-answering1K<n<10K0 likes77 downloads2y agoHugging Face14NNEngine /English-Hindi_Translation 📘 README.md 👉 Copy everything below into your repository README.md English–Hindi Massive Synthetic Translation Dataset 🧠 Overview This dataset is a large-scale synthetic parallel corpus for English → Hindi machine translation, designed to stress-test modern sequence-to-sequence models, tokenizers, and large-scale training pipelines. The corpus contains 10 million aligned sentence pairs generated using a high-entropy template engine with: 100+ subjects 100+… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/English-Hindi_Translation.texttranslation10M<n<100M0 likes76 downloads8mo agoHugging Face15vikasaivyas /hindi-novel-sft-dataset 📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस) यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है। 🌟 प्रमुख विशेषताएँ (Key Highlights) 10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ। 100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.texttext-generationn<1K0 likes73 downloads17d agoHugging Face16suyash2739 /Hinglish About Cleaned dataset from cmu_hinglish_dog [https://huggingface.co/datasets/cmu_hinglish_dog ] texttranslation1K<n<10K2 likes69 downloads2y agoHugging Face17ameykaran /hindi-text-corpus Hindi text corpus Gathered and cleaned from IndicCorpV2 Hindi corpus. text1M<n<10M0 likes69 downloads1y agoHugging Face18Sheeba2026 /bharatvani-hindi-showcase BharatVani Hindi Speech Corpus • Public Interactive Showcase 150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents Curated & Mastered by BharatVani AI • TheCreatorOS 1. Interactive Dataset Preview This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips). Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.audiotext-to-speechn<1K0 likes63 downloads6d agoHugging Face19AIhnIndicRag /Hindi_Fiqatexttext-retrieval10K<n<100K0 likes61 downloads2y agoHugging Face20Yugrathee28 /Hinglish-dataset 🇮🇳 Hinglish Dataset — 1.4 Million Samples Industrial-Grade Code-Mixed NLP Dataset | By ScaleIndia AI · Founder: Yug Rathee This repository contains a 5,000-row teaser sample from the full 1.46 Million+ Hinglish comment dataset built by Scaling YUG (Founder: Yug Rathee(yugrathee28@gmail.com)). Provided strictly for research and evaluation purposes only. Commercial use, redistribution, or production-model training requires explicit written consent from… See the full description on the dataset page: https://huggingface.co/datasets/Yugrathee28/Hinglish-dataset.tabular1K<n<10K2 likes59 downloads5mo agoHugging Face21bingbangboom /adaption-hinglish-transliterate-dataset Adaption Hinglish Transliterate Dataset Dataset Description This dataset contains 77,471 pairs of raw Hindi text captured via Automatic Speech Recognition (ASR) in Devanagari script and their corresponding clean transliterations into Romanized Hinglish. The samples demonstrate the correction of ASR artifacts and the application of Anglicized Hinglish conventions while preserving the original meaning. Each entry consists of an original system prompt instructing… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/adaption-hinglish-transliterate-dataset.texttranslation10K<n<100K1 likes58 downloads3mo agoHugging Face22AIhnIndicRag /Hindi_Climate_Fevertexttext-retrieval1M<n<10M0 likes56 downloads2y agoHugging Face23AmareshHebbar /hindi-medical-sft Hindi Medical Reasoning (Medical-o1-SFT) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Medical questions → detailed chain-of-thought reasoning and clinical answers Why download this Fine-tune models for Hindi-language medical Q&A, build ABDM-compatible clinical assistants, or create multilingual medical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/hindi-medical-sft.texttext-generation10K<n<100K0 likes54 downloads3mo agoHugging Face24OdiaGenAI /instruction_set_hindi_1035The dataset has been created using OliveFarm web application. Following domains have been covered in this dataset:- Art Sports (Cricket, Football, Olympics) Politics History Cooking Environment Music Contributors: - Shahid Parul. textquestion-answering1K<n<10K1 likes52 downloads3y agoHugging Face25ar5entum /hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below: https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file https://github.com/piyushmakhija5/hinglishNorm https://github.com/ishan00/translation-for-code-switching-acl/tree/master text100K<n<1M0 likes51 downloads2y agoHugging Face26uvaidya /hindi-mc4-processedtext1M<n<10M0 likes51 downloads8mo agoHugging Face27ketav /parakeet-hindi-asr Parakeet Hindi-English Bilingual ASR Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition. Quick Start # Download pip install huggingface_hub huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr # Install dependencies pip install nemo_toolkit[asr] bitsandbytes sentencepiece # Train (after updating paths in config) cd parakeet-hindi-asr/scripts python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.textautomatic-speech-recognition100K<n<1M0 likes49 downloads9mo agoHugging Face28Nebulixlabs /Saraswati-Hindi Saraswati-Hindi Saraswati-Hindi is an English-to-Hindi parallel text dataset containing automatically translated English sentences and their corresponding Hindi translations. The dataset was created using the MyMemory Translation API to translate English text into Hindi (en → hi). It is intended for research, experimentation, and development of English-to-Hindi natural language processing (NLP) and machine translation systems. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Nebulixlabs/Saraswati-Hindi.texttext-generationn<1K0 likes48 downloads29d agoHugging Face29BhabhaAI /hindi-RAG-20ktext10K<n<100K3 likes46 downloads3y agoHugging Face30me-nabi /hindikrishi-farmer-advisory-dataset 🌾 HindiKrishi — Farmer Advisory Dataset 21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines. Dataset Details Detail Value Total Examples 21,069 Languages Hindi (primary), English Format JSONL (instruction, input, output) Domain Indian agriculture — crop diseases, pesticides, fertilizers, schemes License Apache 2.0 Format Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.texttext-generation10K<n<100K0 likes45 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.