CoolFace
31 results

hin

manavtabbly /hindi_audio_dataset_testaudion<1K0 likes4.4k downloads11mo agoHugging FaceKathirKs /fineweb-edu-hindi Fineweb-edu-hindi Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2. The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer. Hardware Resources: The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process. Code: Github: fineweb-translation Contact: If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.text100M<n<1B8 likes4.3k downloads2y agoHugging Facezicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes4.1k downloads3y agoHugging Facehudsonburke /rat-hindlimb-mocap Rat Hindlimb Motion Capture Data Processed motion capture data from rat hindlimb gait analysis experiments. Dataset Structure processed/ ├── {subject_id}/ │ ├── markers.parquet # Marker positions (long format) │ ├── forceplates.parquet # Force plate data (long format) │ ├── events.parquet # Gait events (foot strike/off) │ └── sessions.parquet # Per-session anthropometrics └── ... Files markers.parquet: Time series of… See the full description on the dataset page: https://huggingface.co/datasets/hudsonburke/rat-hindlimb-mocap.tabular1B<n<10B1 likes3.9k downloads1mo agoHugging Faceagarwalayushi /hinglish Hinglish Concatenated Audio Dataset A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema. At a Glance Stat Value Total clips 815,171 Total Estimated Hours 2,264+ Unique speakers 6,304 Raw audio size ~243 GB Languages Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.audioautomatic-speech-recognition100K<n<1M7 likes2.1k downloads5mo agoHugging FaceAdaMLLab /HinMix HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filtering, and performs cross-dataset deduplication. We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that HinMix outperforms the previous state-of-the-art, CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/HinMix.texttext-generation100M<n<1B1 likes1.7k downloads8mo agoHugging Face