datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romanized_hindi
Romanized Hindi Dataset
Dataset Description
The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation.
It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods.
The dataset is designed for training and evaluating Hindi↔Roman transliteration models.
Language(s): Hindi, Romanized Hindi
Size: ~1.82M rows
License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.SQuAD_HindiThis dataset is created by translating a part of the Stanford QA dataset.
It contains 5k QA pairs from the original SQuad dataset translated to Hindi using the googletrans api.
hatecheck-hindi
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-hindi.hindi-english-code-mixed-tweets-sentimenthindi-article-summarization
Summary
hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.Hinglish-Hindi-Transliteration-Dataset
Hinglish-Hindi Transliteration Dataset
We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.Movie_Review_Sentiment_HindiHindi_sentimentmindbridge-phq9-hindi-seeds
MindBridge Hindi PHQ-9/GAD-7 — Gold Seeds (144 rows)
Hand-authored Hindi seeds for PHQ-9 + GAD-7 screening across three personas
(postnatal_mother, older_woman, man) in 1:1:1 distribution. Authored via
SuperWhisper Scribe with cloud LLM post-process; all rows
human-reviewed with review_status=accepted.
This seed set drives Phase B teacher expansion (in-context exemplars for
Gemma 4 26B-A4B MoE on Vertex MaaS) plus 24 Item-9 (suicidality) extras
authored separately. See companion… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-seeds.Hindi-Non-STEM-QA-MCQ-DatasetDataset Description:
This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains.
The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.HindiHindi_Train_ClosedDomainQAThe dataset is the Hindi-only and processed version of
https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.hi
https://huggingface.co/datasets/xtreme
https://huggingface.co/datasets/xquad
https://huggingface.co/datasets/databricks/databricks-dolly-15k/viewer/default/train?p=17&f[category][value]=%27closed_qa%27 (closed-qa only)
Hindi-Speech-Dataset
🎧 Hindi Speech Dataset
The Hindi Speech Dataset is a high-quality and structured speech audio dataset developed to support modern AI systems that rely on diverse audio data and scalable voice data. It contains 132 hours of recordings distributed across 565 files, available in MP3 and WAV formats, with a total size of 101 MB. This carefully curated audio dataset provides balanced speaker representation with 49% female and 51% male contributors, covering an age range from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hindi-Speech-Dataset.function-calling-dataset-Hindi-englishhindi-speech-recognition-dataset
Hindi Telephone Dialogues Dataset - 760 Hours
Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.hindi-headline-article-generation
Summary
hindi-headline-article-generation is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation.Hindi-Poetry-Dataset
Hindi Transliteration of Urdu Poetry Dataset
Welcome to the Hindi Transliteration of Urdu Poetry Dataset! This dataset features Hindi transliterations of traditional Urdu poetry. Each entry in the dataset includes two columns:
Title: The transliterated title of the poem in Hindi.
Poem: The transliterated text of the Urdu poem rendered in Hindi script.
This dataset is perfect for researchers and developers working on cross-script language processing, transliteration models, and… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Hindi-Poetry-Dataset.hindi-speech-recognition-dataset
Hindi Speech Dataset for recognition task
Dataset comprises 760 hours of telephone dialogues in Hindi, collected from 1,000+ native speakers across various topics and domains. This dataset boasts an impressive 95% sentence accuracy rate, making it a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/hindi-speech-recognition-dataset.hindi_visual_genomeenglish-hindi-vocab-flashcardsHindi_Mithai
Dataset Card for Indian Sweets
hin_dialect_classificationHindiNewSummaries
How to use this dataset
# You can load data using following code and then split into train and validation set
from datasets import load_dataset
data = load_dataset("d0r1h/HindiNewSummaries")
Licensing
Note:
This license applies to the dataset curation and summaries. Original article copyrights remain with their publishers.
This dataset contains news articles and summaries collected from publicly
available news websites.
The original copyright of the article text… See the full description on the dataset page: https://huggingface.co/datasets/d0r1h/HindiNewSummaries.hindi-sentimentslabel: {'Neutral': 0, 'Positive': 1, 'Negative': 2}
inshorts-hindienglish-hindi-colloquial-datasetA curated dataset of colloquial English phrases and their corresponding Hindi translations. This dataset focuses on informal language, including slang, idioms, and everyday expressions, making it ideal for training models that handle casual conversations.
Dataset Details:
Size:e.g., 500+ phrase pairs]
Source: Collected from publicly available conversational datasets, social media, and crowdsourced contributions.
Language Pair: English → Hindi
Annotations: Each phrase pair is manually verified… See the full description on the dataset page: https://huggingface.co/datasets/bajpaideeksha/english-hindi-colloquial-dataset.hindi_instruction_set_187Kbengali-hindi-number-blindspot
Blind Spots of Frontier Models: Bengali & Hindi Number Word-to-Digit Conversion
Summary
This dataset documents a critical blind spot in small open-source language models: failure to correctly convert Bengali and Hindi number words into their digit equivalents. Bengali and Hindi share the South Asian number system (hazar/হাজার, lakh/লাখ, crore/কোটি), and all three tested models consistently fail at this fundamental conversion step.
IMPORTANT: Arithmetic calculation errors… See the full description on the dataset page: https://huggingface.co/datasets/wong132/bengali-hindi-number-blindspot.sentence_retrieval_hindi_SFTe5_FT_sentence_retrieval_task_Hindi
