datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi_audio_dataset_testfineweb-edu-hindi
Fineweb-edu-hindi
Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2.
The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer.
Hardware Resources:
The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process.
Code:
Github: fineweb-translation
Contact:
If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
rat-hindlimb-mocap
Rat Hindlimb Motion Capture Data
Processed motion capture data from rat hindlimb gait analysis experiments.
Dataset Structure
processed/
├── {subject_id}/
│ ├── markers.parquet # Marker positions (long format)
│ ├── forceplates.parquet # Force plate data (long format)
│ ├── events.parquet # Gait events (foot strike/off)
│ └── sessions.parquet # Per-session anthropometrics
└── ...
Files
markers.parquet: Time series of… See the full description on the dataset page: https://huggingface.co/datasets/hudsonburke/rat-hindlimb-mocap.hinglish
Hinglish Concatenated Audio Dataset
A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema.
At a Glance
Stat
Value
Total clips
815,171
Total Estimated Hours
2,264+
Unique speakers
6,304
Raw audio size
~243 GB
Languages
Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.HinMix
HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filtering, and performs cross-dataset deduplication.
We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that HinMix outperforms the previous state-of-the-art, CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/HinMix.iitb-english-hindi
IITB-English-Hindi Parallel Corpus
About
The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task since 2016 the Hindi-to-English and English-to-Hindi… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/iitb-english-hindi.IndicTTS-Hindi
Hindi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Hindi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.C4-Hindi-Cleaned
Dataset Card for "C4-Hindi-Cleaned"
More Information needed
hindi-english-raw-text-corpus-uncleanedhindi_data_975_originalNews_Hinglish_English
News_Hinglish_English — An English ↔ Hinglish Parallel Corpus
A curated parallel corpus of news-domain text in Hinglish (romanized Hindi-English code-mixed register) paired with corresponding standard English versions. Built to train and evaluate English → Hinglish translation models where existing resources (mostly conversational, e.g., CMU Hinglish DoG) don't cover the news register.
DOI: 10.57967/hf/5120 · License: Apache 2.0 · Downloads: 2,500+
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/suyash2739/News_Hinglish_English.mC4-Hindi-Cleaned
Dataset Card for "mC4-Hindi-Cleaned"
More Information needed
hindi_books_collectionhindi-aggregatedhindawi-journals-2007-2023
Hindawi Academic Papers Dataset (CC BY 4.0 Compatible)
Dataset Description
This dataset contains 299,316 academic research papers from Hindawi Publishing Corporation, carefully filtered to include only papers with licenses compatible with CC BY 4.0. The dataset includes comprehensive metadata for each paper including titles, authors, journal information, publication years, DOIs, and full-text content.
Dataset Summary
Total Papers: 299,316 (filtered from 299… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/hindawi-journals-2007-2023.syspin_hindi_mergedwhisper-hindi-preprocessedHindi-1482HrsIndicVoices-R_Hindibaarat-hindi-pretrain-dataMUCS-Hinglish
MUCS
Dataset Description
This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset.
This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2.
As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here.
In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.HINMIX_hi-en
Dataset Card for Hindi English Codemix Dataset - HINMIX
HINMIX is a massive parallel codemixed dataset for Hindi-English code switching.
See the 📚 paper on arxiv to dive deep into this synthetic codemix data generation pipeline.
Dataset contains 4.2M fully parallel sentences in 6 Hindi-English forms.
Further, we release gold standard codemix dev and test set manually translated by proficient bilingual annotators.
Dev Set consists of 280 examples
Test set consists of 2507 examples… See the full description on the dataset page: https://huggingface.co/datasets/kartikagg98/HINMIX_hi-en.hindi-ocr
Dataset Card for Dataset Name
Line-level text and images from https://heidata.uni-heidelberg.de/dataset.xhtml?persistentId=doi:10.11588/data/EGOKEI
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/apjanco/hindi-ocr.MMLU_HindiHindi version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2hindustani-raag-small
hindustani-raag-small
50 raags of Hindustani classical music, chunked from YouTube recordings, each
chunk carrying the hand-annotated tonic (Sa) of the recording it came from.
raags
50
train clips
1810
test clips
150
clip length
20-60 s
source videos
412
Columns
audio — an mp3 chunk, 20-60 s, cut from a full recording.
label — ClassLabel over the 50 raag names.
tonic_hz — the recording's Sa in Hz, annotated by ear and snapped to a peak… See the full description on the dataset page: https://huggingface.co/datasets/neerajaabhyankar/hindustani-raag-small.hindi_karya_mergedhda_nli_hindiThis dataset is a recasted version of the Hindi Discourse Analysis Dataset used to train models for Natural Language Inference Tasks in Low-Resource Languages like Hindi.S2R_Shrutilipi_hindi
Paytmlabs/S2R_Shrutilipi_hindi
Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training.
Viewing samples on Hugging Face
The hindi config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows.
To inspect examples in the browser, open the Subset (config) drop-down and choose hindi_text_samples — text and continuation only (~2000 rows).
Ultravox training should keep using subset hindi (full audio).… See the full description on the dataset page: https://huggingface.co/datasets/Paytmlabs/S2R_Shrutilipi_hindi.
