hin
Datasets
All datasets matching “hin”hindi_audio_dataset_testfineweb-edu-hindi
Fineweb-edu-hindi
Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2.
The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer.
Hardware Resources:
The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process.
Code:
Github: fineweb-translation
Contact:
If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
rat-hindlimb-mocap
Rat Hindlimb Motion Capture Data
Processed motion capture data from rat hindlimb gait analysis experiments.
Dataset Structure
processed/
├── {subject_id}/
│ ├── markers.parquet # Marker positions (long format)
│ ├── forceplates.parquet # Force plate data (long format)
│ ├── events.parquet # Gait events (foot strike/off)
│ └── sessions.parquet # Per-session anthropometrics
└── ...
Files
markers.parquet: Time series of… See the full description on the dataset page: https://huggingface.co/datasets/hudsonburke/rat-hindlimb-mocap.hinglish
Hinglish Concatenated Audio Dataset
A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema.
At a Glance
Stat
Value
Total clips
815,171
Total Estimated Hours
2,264+
Unique speakers
6,304
Raw audio size
~243 GB
Languages
Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.HinMix
HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filtering, and performs cross-dataset deduplication.
We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that HinMix outperforms the previous state-of-the-art, CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/HinMix.
