CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pfin123 /hindi-aggregatedtext100K<n<1M2 likes649 downloads4y agoHugging Face02Sheeba2026 /bharatvani-hindi-speech-corpusgated BharatVani Hindi Speech Corpus (150-Hour Studio Dataset) Proprietary Speech Asset • TheCreatorOS • BharatVani AI 1. Overview The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi. Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.audiotext-to-speech100K<n<1M1 likes285 downloads6d agoHugging Face03QCRI /LlamaLens-Hindi LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi.texttext-classification100K<n<1M1 likes117 downloads2y agoHugging Face04FreedomIntelligence /alpaca-gpt4-hindiThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K1 likes115 downloads3y agoHugging Face05FreedomIntelligence /evol-instruct-hindiThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K2 likes111 downloads3y agoHugging Face06QCRI /LlamaLens-Hindi-Native LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi-Native.texttext-classification100K<n<1M0 likes105 downloads2y agoHugging Face07nirantk /chaii-hindi-and-tamil-question-answeringtextquestion-answering1K<n<10K0 likes86 downloads3y agoHugging Face08AIhnIndicRag /Hindi_Fevertexttext-retrieval1M<n<10M0 likes86 downloads2y agoHugging Face09OdiaGenAI /sentiment_analysis_hindiConventions followed to decide the polarity: - labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'. labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.texttext-classification1K<n<10K2 likes80 downloads3y agoHugging Face10rmahesh /UP_CET_Hindi_examstextquestion-answering1K<n<10K0 likes77 downloads2y agoHugging Face11NNEngine /English-Hindi_Translation 📘 README.md 👉 Copy everything below into your repository README.md English–Hindi Massive Synthetic Translation Dataset 🧠 Overview This dataset is a large-scale synthetic parallel corpus for English → Hindi machine translation, designed to stress-test modern sequence-to-sequence models, tokenizers, and large-scale training pipelines. The corpus contains 10 million aligned sentence pairs generated using a high-entropy template engine with: 100+ subjects 100+… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/English-Hindi_Translation.texttranslation10M<n<100M0 likes76 downloads8mo agoHugging Face12vikasaivyas /hindi-novel-sft-dataset 📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस) यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है। 🌟 प्रमुख विशेषताएँ (Key Highlights) 10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ। 100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.texttext-generationn<1K0 likes73 downloads17d agoHugging Face13ameykaran /hindi-text-corpus Hindi text corpus Gathered and cleaned from IndicCorpV2 Hindi corpus. text1M<n<10M0 likes69 downloads1y agoHugging Face14Sheeba2026 /bharatvani-hindi-showcase BharatVani Hindi Speech Corpus • Public Interactive Showcase 150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents Curated & Mastered by BharatVani AI • TheCreatorOS 1. Interactive Dataset Preview This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips). Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.audiotext-to-speechn<1K0 likes63 downloads7d agoHugging Face15AIhnIndicRag /Hindi_Fiqatexttext-retrieval10K<n<100K0 likes61 downloads2y agoHugging Face16AIhnIndicRag /Hindi_Climate_Fevertexttext-retrieval1M<n<10M0 likes56 downloads2y agoHugging Face17AmareshHebbar /hindi-medical-sft Hindi Medical Reasoning (Medical-o1-SFT) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Medical questions → detailed chain-of-thought reasoning and clinical answers Why download this Fine-tune models for Hindi-language medical Q&A, build ABDM-compatible clinical assistants, or create multilingual medical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/hindi-medical-sft.texttext-generation10K<n<100K0 likes54 downloads3mo agoHugging Face18OdiaGenAI /instruction_set_hindi_1035The dataset has been created using OliveFarm web application. Following domains have been covered in this dataset:- Art Sports (Cricket, Football, Olympics) Politics History Cooking Environment Music Contributors: - Shahid Parul. textquestion-answering1K<n<10K1 likes52 downloads3y agoHugging Face19ar5entum /hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below: https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file https://github.com/piyushmakhija5/hinglishNorm https://github.com/ishan00/translation-for-code-switching-acl/tree/master text100K<n<1M0 likes51 downloads2y agoHugging Face20uvaidya /hindi-mc4-processedtext1M<n<10M0 likes51 downloads8mo agoHugging Face21ketav /parakeet-hindi-asr Parakeet Hindi-English Bilingual ASR Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition. Quick Start # Download pip install huggingface_hub huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr # Install dependencies pip install nemo_toolkit[asr] bitsandbytes sentencepiece # Train (after updating paths in config) cd parakeet-hindi-asr/scripts python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.textautomatic-speech-recognition100K<n<1M0 likes49 downloads9mo agoHugging Face22Nebulixlabs /Saraswati-Hindi Saraswati-Hindi Saraswati-Hindi is an English-to-Hindi parallel text dataset containing automatically translated English sentences and their corresponding Hindi translations. The dataset was created using the MyMemory Translation API to translate English text into Hindi (en → hi). It is intended for research, experimentation, and development of English-to-Hindi natural language processing (NLP) and machine translation systems. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Nebulixlabs/Saraswati-Hindi.texttext-generationn<1K0 likes48 downloads29d agoHugging Face23BhabhaAI /hindi-RAG-20ktext10K<n<100K3 likes46 downloads3y agoHugging Face24me-nabi /hindikrishi-farmer-advisory-dataset 🌾 HindiKrishi — Farmer Advisory Dataset 21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines. Dataset Details Detail Value Total Examples 21,069 Languages Hindi (primary), English Format JSONL (instruction, input, output) Domain Indian agriculture — crop diseases, pesticides, fertilizers, schemes License Apache 2.0 Format Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.texttext-generation10K<n<100K0 likes45 downloads2mo agoHugging Face25Vidyaapati-Hindi-Konkani /Vidyaapati-Hindi-Konkani-Testsettext1K<n<10K0 likes44 downloads22d agoHugging Face26srajwal1 /hindi-driving-questionsThis is part of Global Exam project by C4AI Also, the dataset has been checked against the script given here (https://github.com/for-ai/global-exams/blob/main/any_language/dataset_checker.py) and then uploaded. The dataset in this repository is part of hindi cohort from this link: https://www.transport.telangana.gov.in/html/pdf/hindi.pdf textn<1K0 likes40 downloads2y agoHugging Face27BhabhaAI /alpaca-gpt4-hindi-trans Alpaca GPT4 Hindi This dataset is hindi translated filtered version of alpaca-gpt4 using IndicTrans2. text10K<n<100K0 likes39 downloads3y agoHugging Face28uvaidya /hindi-mc4-20btext1M<n<10M0 likes37 downloads8mo agoHugging Face29BhabhaAI /Cross-Hindi-Hinglish-chat Cross Hindi Hinglish Chat This dataset is a subset of OpenHermes where some part is converted to either Hindi or Hinglish.Note: This is in raw form. You must add "Reply in Hindi", "Reply in English" kind texts where appropriate.row_ids correspond to row id starting from 0 for OpenHermes English dataset. texttext-generation10K<n<100K1 likes35 downloads3y agoHugging Face30uvaidya /hindi-culturax-20btext10M<n<100M0 likes35 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.