datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HinMix
HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filtering, and performs cross-dataset deduplication.
We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that HinMix outperforms the previous state-of-the-art, CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/HinMix.hindawi-journals-2007-2023
Hindawi Academic Papers Dataset (CC BY 4.0 Compatible)
Dataset Description
This dataset contains 299,316 academic research papers from Hindawi Publishing Corporation, carefully filtered to include only papers with licenses compatible with CC BY 4.0. The dataset includes comprehensive metadata for each paper including titles, authors, journal information, publication years, DOIs, and full-text content.
Dataset Summary
Total Papers: 299,316 (filtered from 299… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/hindawi-journals-2007-2023.Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.mC4-hindi
Dataset Card for "mC4-hindi"
This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts.
This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.long_context_hindi
Dataset
This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations.
This dataset contains only Hindi as of now
Information
First this dataset is mainly for long context training
The minimum len is 6000 and maximum len is 3754718
Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources.
Sources:
Hinglish TOP Dataset
CMU English Dog
HinGE
PHINC
source : 1 - Human Annotated ,
source : 0 - Synthetically Generated
Hindawi-Books-dataset
Dataset Card for "Hindawi Books Dataset"
Hindawi Books Dataset is a large collection of more than 3000 books written in Modern Standard Arabic.
Dataset Description
Hindawi Books Dataset offers a rich and diverse collection of literary works, covering various topics and genres, all written in Modern Standard Arabic. The dataset includes information about each book, such as the title, author name, book abstract, and a link to access the complete text online. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/alielfilali01/Hindawi-Books-dataset.OSCAR-2301-Hindi-Cleaned
Dataset Card for "OSCAR-2301-Hindi-Cleaned-2.0"
More Information needed
Hinglish_Dataset_instruction_and_rawtask1009_pib_translation_bengali_hindi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1009_pib_translation_bengali_hindi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1009_pib_translation_bengali_hindi.Qwen2.5-7B-Instruct-Self-Calibration
Efficient Test-Time Scaling via Self-Calibration
This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.task427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.English-Hinglish-TOP
English Hinglish (TOP Dataset)
This dataset is generated from Hinglish-TOP Dataset.
Data distribution:
Train a. Human Generated - 6513 b. Synthetically generated - 170083
Validation a. Human Generated - 1390 b. Synthetically generated - 0
Test a. Human Generated - 6513 b. Synthetically generated - 0
shreyansh-hinglish-english-stem-500k
🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus
Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations.
📖 Overview
In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.hinglish-instruct-dataset
Akshar Hinglish Instruct
Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns.
1. Dataset Overview
Total Examples: 10,378
Base Set: 9,999 instruction-following pairs
Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.hintedselfteacher-nemotron-math-v2-AoPS
hintedselfteacher-nemotron-math-v2-AoPS
This dataset contains a training-ready hinted self-teacher split derived from the AoPS split of nvidia/Nemotron-Math-v2.
The source problems were filtered to the AoPS split with the medium/notool solve rate between 2 and 6. Hints were generated with GPT-5.5 medium using an h17_nt hint-generation prompt. This hint type was close to the best hint type found after doing hint mutations, based on qualitative analysis of token-level hinted… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/hintedselfteacher-nemotron-math-v2-AoPS.mmlu-hint-faithfulness-traces
MMLU and GPQA hint faithfulness traces
Model-specific configurations
Configuration
Rows
Model and content
qwen3-8b
3,486
Existing Qwen3-8B generations and judgments
qwen3-8b_answer_changes
1,137
Existing Qwen3-8B valid answer changes
gpt-5.6-luna
3,486
Luna final responses, observable reasoning summaries and judgments
gpt-5.6-luna_answer_changes
966
Luna valid answer changes with the same labels
All sets contain MMLU and GPQA and use a test… See the full description on the dataset page: https://huggingface.co/datasets/shiv96/mmlu-hint-faithfulness-traces.English-Hinglish
English Hinglish
English to Hinglish Dataset processed from findnitai/english-to-hinglish.
Sources:
Hinglish TOP Dataset
CMU English Dog
HinGE
PHINC
nisram-hindi-text-0.0
Dataset Card for nisram-hindi-text-0.0
Quick Summary
Language: Hindi
Size: ~618,000 raw text samples (206k from each source, before deduplication) and 602,000 after deduplication
Content: Short- to medium-length Hindi text (up to ~5000 characters), drawn from web-crawled corpora
Fields: One field text (string) containing the Hindi text
Sources:
mC4-Hindi-Cleaned-3.0
OSCAR-2301-Hindi-Cleaned-2.0
ai4bharat/sangraha (Hindi, verified)
Deduplication: Performed… See the full description on the dataset page: https://huggingface.co/datasets/nis12ram/nisram-hindi-text-0.0.task1204_atomic_classification_hinderedby
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1204_atomic_classification_hinderedby
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1204_atomic_classification_hinderedby.hindi-novel-sft-dataset
📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस)
यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है।
🌟 प्रमुख विशेषताएँ (Key Highlights)
10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ।
100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.hinglish_self_instruct_v0
Hinglish Instruct Dataset using Self Instruct method
The prompt used for generating the samples:
You are asked to come up with a set of 50 diverse task instructions in Hinglish or Hindi.
These task instructions will be given to a GPT model and we will evaluate the GPT model for completing the instructions.
Here are the requirements:
1. Try not to repeat the verb for each instruction to maximize diversity.
2. The language used for the instruction also should be diverse. For example… See the full description on the dataset page: https://huggingface.co/datasets/smangrul/hinglish_self_instruct_v0.Hindi-Niband
Dataset Name: Hindi- Niband (Massive Hindi language Text Dataset)
Dataset Overview
This dataset is a comprehensive collection of text data consisting of more than 10 billion tokens. It encompasses a wide range of sources, including Wikipedia articles, news articles, email transcripts, and generated prompt text. Specific Hindi language data columns have been extracted from the CulturaX dataset, which is a large, cleaned, and multilingual dataset for large language models.… See the full description on the dataset page: https://huggingface.co/datasets/thinkedgeAI/Hindi-Niband.Hindi-speech-instruct
Hindi LLaMA-Omni Instruct Dataset
A Hindi speech instruction-following dataset designed for training speech-language models such as LLaMA-Omni. Each example pairs a spoken Hindi user question (audio) with a text assistant response.
Dataset Summary
Property
Value
Language
Hindi (hi)
Total examples
~110,718
Train split
~105,000 examples (batches 001–210)
Validation split
~5,500 examples (batches 211–222)
Audio format
FLAC, 16,000 Hz mono… See the full description on the dataset page: https://huggingface.co/datasets/Pastaaaaa2003/Hindi-speech-instruct.hinglish-bench
Hinglish-Bench
📄 Paper: Hinglish-Bench — Reference-Free Benchmark for LLM Hinglish Text
Generation
(gist preprint)
A reference-free benchmark for measuring how well LLMs generate natural
Roman-script Hinglish — the Hindi-English code-mixing that hundreds of
millions of Indians actually speak, type, and read online.
Reference-free by design. Hinglish has no canonical spelling and no single
"correct" rendering, so there are no gold references and no BLEU. Quality is… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/hinglish-bench.hindi-article-summarization
Summary
hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.hindi-medical-sft
Hindi Medical Reasoning (Medical-o1-SFT)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Medical questions → detailed chain-of-thought reasoning and clinical answers
Why download this
Fine-tune models for Hindi-language medical Q&A, build ABDM-compatible clinical assistants, or create multilingual medical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/hindi-medical-sft.input_ablation_qwen3_8b_mmlu_hint
Training Language Models to Explain Their Own Computations (Input Ablations)
This dataset is part of the work presented in the paper "Training Language Models to Explain Their Own Computations".
It specifically contains data for the Input Ablations task for the Qwen3-8B target model. In this task, explainer models are trained to predict how removing "hint" tokens from an MMLU prompt with a hint changes the output of Qwen3-8B. This helps in understanding the causal relationships… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/input_ablation_qwen3_8b_mmlu_hint.qwen3-4b-base-self-distillation-hints
Qwen3-4B-Base Self-Distillation Hints
4,481 unique accepted math problems, each with a verifiable target answer,
a reference worked solution, and an h17_nt private problem-solving hint.
The intended student is Qwen/Qwen3-4B-Base. Hints were generated by
GPT-5.6 Sol (gpt-5.6-sol), medium reasoning effort, default service tier,
not by Qwen. No student rollout traces were supplied to the hint generator.
Intended Use
Use prompt for the student and hint_text as… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/qwen3-4b-base-self-distillation-hints.Saraswati-Hindi
Saraswati-Hindi
Saraswati-Hindi is an English-to-Hindi parallel text dataset containing automatically translated English sentences and their corresponding Hindi translations.
The dataset was created using the MyMemory Translation API to translate English text into Hindi (en → hi). It is intended for research, experimentation, and development of English-to-Hindi natural language processing (NLP) and machine translation systems.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Nebulixlabs/Saraswati-Hindi.
