datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi-aggregatedbharatvani-hindi-speech-corpus
BharatVani Hindi Speech Corpus (150-Hour Studio Dataset)
Proprietary Speech Asset • TheCreatorOS • BharatVani AI
1. Overview
The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi.
Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources.
Sources:
Hinglish TOP Dataset
CMU English Dog
HinGE
PHINC
source : 1 - Human Annotated ,
source : 0 - Synthetically Generated
hindsight-neglect-10shot
inverse-scaling/hindsight-neglect-10shot (‘The Floating Droid’)
General description
This task tests whether language models are able to assess whether a bet was worth taking based on its expected value. The author provides few shot examples in which the model predicts whether a bet is worthwhile by correctly answering yes or no when the expected value of the bet is positive (where the model should respond that ‘yes’, taking the bet is the right decision) or negative (‘no’… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling/hindsight-neglect-10shot.LlamaLens-Hindi
LlamaLens: Specialized Multilingual LLM Dataset
Overview
LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi.
LlamaLens
This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation.
Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi.alpaca-gpt4-hindiThe dataset is used in the research related to MultilingualSIFT.
evol-instruct-hindiThe dataset is used in the research related to MultilingualSIFT.
LlamaLens-Hindi-Native
LlamaLens: Specialized Multilingual LLM Dataset
Overview
LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi.
LlamaLens
This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation.
Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi-Native.hinglish-instruct-dataset
Akshar Hinglish Instruct
Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns.
1. Dataset Overview
Total Examples: 10,378
Base Set: 9,999 instruction-following pairs
Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.chaii-hindi-and-tamil-question-answeringHindi_Feversentiment_analysis_hindiConventions followed to decide the polarity: -
labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos
labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'.
labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.UP_CET_Hindi_examsEnglish-Hindi_Translation
📘 README.md
👉 Copy everything below into your repository README.md
English–Hindi Massive Synthetic Translation Dataset
🧠 Overview
This dataset is a large-scale synthetic parallel corpus for English → Hindi machine translation, designed to stress-test modern sequence-to-sequence models, tokenizers, and large-scale training pipelines.
The corpus contains 10 million aligned sentence pairs generated using a high-entropy template engine with:
100+ subjects
100+… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/English-Hindi_Translation.hindi-novel-sft-dataset
📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस)
यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है।
🌟 प्रमुख विशेषताएँ (Key Highlights)
10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ।
100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.Hinglish
About
Cleaned dataset from cmu_hinglish_dog [https://huggingface.co/datasets/cmu_hinglish_dog ]
hindi-text-corpus
Hindi text corpus
Gathered and cleaned from IndicCorpV2 Hindi corpus.
bharatvani-hindi-showcase
BharatVani Hindi Speech Corpus • Public Interactive Showcase
150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents
Curated & Mastered by BharatVani AI • TheCreatorOS
1. Interactive Dataset Preview
This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips).
Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.Hindi_FiqaHinglish-dataset
🇮🇳 Hinglish Dataset — 1.4 Million Samples
Industrial-Grade Code-Mixed NLP Dataset | By ScaleIndia AI · Founder: Yug Rathee
This repository contains a 5,000-row teaser sample from the full 1.46 Million+ Hinglish comment dataset built by Scaling YUG (Founder: Yug Rathee(yugrathee28@gmail.com)).
Provided strictly for research and evaluation purposes only.
Commercial use, redistribution, or production-model training requires explicit written consent from… See the full description on the dataset page: https://huggingface.co/datasets/Yugrathee28/Hinglish-dataset.adaption-hinglish-transliterate-dataset
Adaption Hinglish Transliterate Dataset
Dataset Description
This dataset contains 77,471 pairs of raw Hindi text captured via Automatic Speech Recognition (ASR) in Devanagari script and their corresponding clean transliterations into Romanized Hinglish. The samples demonstrate the correction of ASR artifacts and the application of Anglicized Hinglish conventions while preserving the original meaning.
Each entry consists of an original system prompt instructing… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/adaption-hinglish-transliterate-dataset.Hindi_Climate_Feverhindi-medical-sft
Hindi Medical Reasoning (Medical-o1-SFT)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Medical questions → detailed chain-of-thought reasoning and clinical answers
Why download this
Fine-tune models for Hindi-language medical Q&A, build ABDM-compatible clinical assistants, or create multilingual medical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/hindi-medical-sft.instruction_set_hindi_1035The dataset has been created using OliveFarm web application.
Following domains have been covered in this dataset:-
Art
Sports (Cricket, Football, Olympics)
Politics
History
Cooking
Environment
Music
Contributors: -
Shahid
Parul.
hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below:
https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file
https://github.com/piyushmakhija5/hinglishNorm
https://github.com/ishan00/translation-for-code-switching-acl/tree/master
hindi-mc4-processedparakeet-hindi-asr
Parakeet Hindi-English Bilingual ASR
Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition.
Quick Start
# Download
pip install huggingface_hub
huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr
# Install dependencies
pip install nemo_toolkit[asr] bitsandbytes sentencepiece
# Train (after updating paths in config)
cd parakeet-hindi-asr/scripts
python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.Saraswati-Hindi
Saraswati-Hindi
Saraswati-Hindi is an English-to-Hindi parallel text dataset containing automatically translated English sentences and their corresponding Hindi translations.
The dataset was created using the MyMemory Translation API to translate English text into Hindi (en → hi). It is intended for research, experimentation, and development of English-to-Hindi natural language processing (NLP) and machine translation systems.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Nebulixlabs/Saraswati-Hindi.hindi-RAG-20khindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.
