datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.Hindi-speech-instruct
Hindi LLaMA-Omni Instruct Dataset
A Hindi speech instruction-following dataset designed for training speech-language models such as LLaMA-Omni. Each example pairs a spoken Hindi user question (audio) with a text assistant response.
Dataset Summary
Property
Value
Language
Hindi (hi)
Total examples
~110,718
Train split
~105,000 examples (batches 001–210)
Validation split
~5,500 examples (batches 211–222)
Audio format
FLAC, 16,000 Hz mono… See the full description on the dataset page: https://huggingface.co/datasets/Pastaaaaa2003/Hindi-speech-instruct.tajik-spoken-instructions
Tajik Spoken Instructions (Q&A)
Synthetic Tajik speech of dictionary and language-exercise questions, each paired with its
written answer — spoken instruction in, text answer out.
648,842 clips · ~540 hours · 2 voices (male + female)
Questions cover word meanings, antonyms, etymology, usage and grammar
Numbers expanded to spoken Tajik
Deduplicated; Latin-script and mis-encoded rows removed
Contents
file
what
audio_*.tar
the wav files
manifest.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-spoken-instructions.instructs2s-snac
InstructS2S-SNAC
Pre-processed dataset for Speech-to-Speech model training.
Original Dataset
This dataset is derived from ICTNLP/InstructS2S-200K.
Processing
The audio has been pre-processed into:
Whisper features: Input audio encoded with Whisper large-v3 encoder
SNAC tokens: Output audio tokenized with SNAC codec (24kHz)
Statistics
Total samples: 23,992
Input audio: ~100h (Whisper features)
Output audio: ~61h (SNAC tokens)
File size: ~86GB… See the full description on the dataset page: https://huggingface.co/datasets/marcosremar2/instructs2s-snac.
