Hindi
Datasets
All datasets matching “Hindi”hindi_audio_dataset_testfineweb-edu-hindi
Fineweb-edu-hindi
Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2.
The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer.
Hardware Resources:
The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process.
Code:
Github: fineweb-translation
Contact:
If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
iitb-english-hindi
IITB-English-Hindi Parallel Corpus
About
The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task since 2016 the Hindi-to-English and English-to-Hindi… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/iitb-english-hindi.IndicTTS-Hindi
Hindi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Hindi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.C4-Hindi-Cleaned
Dataset Card for "C4-Hindi-Cleaned"
More Information needed
