CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Vikhrmodels /Ficbook-Audio-Instruct-10K Ficbook Audio Instruct 10K Synthetic audio instruction dataset for training Russian audio-language models. Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks. Dataset Description This dataset was created for training and evaluating audio-language models on Russian fiction content. Each sample contains: Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model Text: Original text from ficbook stories Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.audioautomatic-speech-recognition1K<n<10K0 likes81 downloads9mo agoHugging Face02Pastaaaaa2003 /Hindi-speech-instructgated Hindi LLaMA-Omni Instruct Dataset A Hindi speech instruction-following dataset designed for training speech-language models such as LLaMA-Omni. Each example pairs a spoken Hindi user question (audio) with a text assistant response. Dataset Summary Property Value Language Hindi (hi) Total examples ~110,718 Train split ~105,000 examples (batches 001–210) Validation split ~5,500 examples (batches 211–222) Audio format FLAC, 16,000 Hz mono… See the full description on the dataset page: https://huggingface.co/datasets/Pastaaaaa2003/Hindi-speech-instruct.audioautomatic-speech-recognition100K<n<1M0 likes55 downloads3mo agoHugging Face03Tohirju /tajik-spoken-instructionsgated Tajik Spoken Instructions (Q&A) Synthetic Tajik speech of dictionary and language-exercise questions, each paired with its written answer — spoken instruction in, text answer out. 648,842 clips · ~540 hours · 2 voices (male + female) Questions cover word meanings, antonyms, etymology, usage and grammar Numbers expanded to spoken Tajik Deduplicated; Latin-script and mis-encoded rows removed Contents file what audio_*.tar the wav files manifest.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-spoken-instructions.text-to-speech100K<n<1M0 likes38 downloads14d agoHugging Face04marcosremar2 /instructs2s-snac InstructS2S-SNAC Pre-processed dataset for Speech-to-Speech model training. Original Dataset This dataset is derived from ICTNLP/InstructS2S-200K. Processing The audio has been pre-processed into: Whisper features: Input audio encoded with Whisper large-v3 encoder SNAC tokens: Output audio tokenized with SNAC codec (24kHz) Statistics Total samples: 23,992 Input audio: ~100h (Whisper features) Output audio: ~61h (SNAC tokens) File size: ~86GB… See the full description on the dataset page: https://huggingface.co/datasets/marcosremar2/instructs2s-snac.text-to-speech10K<n<100K0 likes5 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.