CoolFace
Datasetpublicgated

Pastaaaaa2003/Hindi-speech-instruct

Hindi LLaMA-Omni Instruct Dataset A Hindi speech instruction-following dataset designed for training speech-language models such as LLaMA-Omni. Each example pairs a spoken Hindi user question (audio) with a text assistant response. Dataset Summary Property Value Language Hindi (hi) Total examples ~110,718 Train split ~105,000 examples (batches 001–210) Validation split ~5,500 examples (batches 211–222) Audio format FLAC, 16,000 Hz mono… See the full description on the dataset page: https://huggingface.co/datasets/Pastaaaaa2003/Hindi-speech-instruct.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes55downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.

Pastaaaaa2003/Hindi-speech-instruct · CoolFace