datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi_audio_dataset_testfineweb-edu-hindi
Fineweb-edu-hindi
Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2.
The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer.
Hardware Resources:
The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process.
Code:
Github: fineweb-translation
Contact:
If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
iitb-english-hindi
IITB-English-Hindi Parallel Corpus
About
The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task since 2016 the Hindi-to-English and English-to-Hindi… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/iitb-english-hindi.IndicTTS-Hindi
Hindi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Hindi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.C4-Hindi-Cleaned
Dataset Card for "C4-Hindi-Cleaned"
More Information needed
hindi-english-raw-text-corpus-uncleanedsyspin_hindi_mergedhindi-aggregatedmC4-Hindi-Cleaned
Dataset Card for "mC4-Hindi-Cleaned"
More Information needed
IndicVoices-R_HindiHindi-1482Hrshindi-ocr
Dataset Card for Dataset Name
Line-level text and images from https://heidata.uni-heidelberg.de/dataset.xhtml?persistentId=doi:10.11588/data/EGOKEI
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/apjanco/hindi-ocr.hindi_karya_mergedbaarat-hindi-pretrain-dataS2R_Shrutilipi_hindi
Paytmlabs/S2R_Shrutilipi_hindi
Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training.
Viewing samples on Hugging Face
The hindi config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows.
To inspect examples in the browser, open the Subset (config) drop-down and choose hindi_text_samples — text and continuation only (~2000 rows).
Ultravox training should keep using subset hindi (full audio).… See the full description on the dataset page: https://huggingface.co/datasets/Paytmlabs/S2R_Shrutilipi_hindi.hindi-asr-wdsmC4-hindi
Dataset Card for "mC4-hindi"
This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts.
This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.hindi-gov-vqa_beir This is a copy of https://huggingface.co/datasets/jinaai/hindi-gov-vqa reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hindi-gov-vqa_beir.hindi_indic_voice_rbharatvani-hindi-speech-corpus
BharatVani Hindi Speech Corpus (150-Hour Studio Dataset)
Proprietary Speech Asset • TheCreatorOS • BharatVani AI
1. Overview
The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi.
Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.All_Hindi_ASR_v1.1IIIT-INDIC-HW-WORDS-Hindi
IIIT-INDIC-HW-WORDS-Hindi
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Hindi.baarat-batched-hindi-pre-trainingAll_Hindi_ASR_v1.2OCR-Bench1000-Hindi
OCR-Bench1000-Hindi
1000 synthetic printed-text line images with ground-truth transcriptions,
sampled from a larger locally-held Hindi OCR training corpus.
This is a benchmark/sample release, not the full training set.
Data fields
Field
Description
file_name
relative path to the image (images/...)
text
ground-truth transcription
category
hindi_only / english_only / mixed / numeric_and_symbols
length_bucket
short / medium / long, by character count… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Hindi.long_context_hindi
Dataset
This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations.
This dataset contains only Hindi as of now
Information
First this dataset is mainly for long context training
The minimum len is 6000 and maximum len is 3754718
Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.Hindi_corpus-4096-packed-qtk-1.37M
fhai50032/Hindi_corpus-4096-packed-qtk-1.37M
Packed pretraining corpus, 1,369,441 rows x 4096 tokens = 5.609B tokens, tokenized with
fhai50032/QTK-81K.
Format
column
type
notes
label
list<int32>
exactly 4096 tokens, no padding
raw_label
string
label decoded back to text (redundant, for inspection)
There is no attention_mask column: the corpus is packed, so every position is a real token and
the mask would be all ones on every row.
How… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/Hindi_corpus-4096-packed-qtk-1.37M.All_Hindi_ASR_v1.1HindiDownloadedData2
