CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes4.7k downloads7mo agoHugging Face02ErfanRou /callcc-2k-hours-train-sonioxgated ErfanRou/callcc-2k Persian call-centre ASR silver set (2,000 audio-hours, 2,288 h shipped): the ErfanRou/callcc-ft150-soniox schema plus two auxiliary columns. Machine-labelled, not ground truth. Built by callcc-silver-5k/soniox_pipeline.py from markmuller/call-center-prod-data: 8 kHz stereo telephony split client-side into agent = channel 0 and customer = channel 1, each channel transcribed separately by Soniox stt-async-v5, words packed into 5-28 s windows (15 s mean, natural… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/callcc-2k-hours-train-soniox.audio100K<n<1M0 likes273 downloads7d agoHugging Face03kamilakesbi /cv_for_spd_fr_2k_augmentedaudio1K<n<10K0 likes68 downloads2y agoHugging Face04kamilakesbi /cv_for_spd_fr_augmented_2kaudio1K<n<10K0 likes65 downloads2y agoHugging Face05smutuvi /ndizi-1_sample_2kaudio1K<n<10K0 likes54 downloads1y agoHugging Face06TuNguyen1101 /kaito_2kaudio1K<n<10K1 likes51 downloads4mo agoHugging Face07manishkumar2101114 /all-12-voices-2kaudio1K<n<10K0 likes50 downloads6mo agoHugging Face08twangodev /radiotalk-voices-2k radiotalk-voices-2k 2,000 English reference voices — one 12–30s clip per speaker, selected as the longest qualifying utterance per speaker from LibriTTS-R. Built for zero-shot TTS voice cloning in the radiotalk pipeline. Stats 2,000 voices · 12.03 hours total Duration: min 12.0s · median 21.8s · mean 21.7s · max 30.0s 24 kHz, mono, FLAC-encoded Schema Column Type Description voice_id string Stable 12-hex-char id, derived from (source… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-voices-2k.audiotext-to-speech1K<n<10K0 likes43 downloads5mo agoHugging Face09kamilakesbi /cv_for_spd_fr_2k_std_0.5audio1K<n<10K0 likes37 downloads2y agoHugging Face10axolotl-ai-co /text-vision-audio-2k-testA 2k sample dataset for testing multimodal (text+vision+audio) format. This is compatible with HF's processor apply_chat_template. Load in Axolotl via: datasets: - path: Nanobit/text-vision-audio-2k-test type: chat_template Make sure to download the image and audio via: wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/African_elephant.jpg wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/En-us-African_elephant.oga… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-audio-2k-test.audio1K<n<10K0 likes36 downloads1y agoHugging Face11axolotl-ai-co /text-audio-2k-testA 2k sample dataset for testing multimodal (text+audio) format. This is compatible with HF's processor apply_chat_template. Load in Axolotl via: datasets: - path: Nanobit/text-audio-2k-test type: chat_template Make sure to download the audio via: wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/En-us-African_elephant.oga Audio source: https://upload.wikimedia.org/wikipedia/commons/a/ad/En-us-African_elephant.oga Each sample has the following format… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-audio-2k-test.audio1K<n<10K0 likes36 downloads1y agoHugging Face12mavihsrr /Hindi_TTS_M-2kaudio1K<n<10K0 likes31 downloads2y agoHugging Face13kamilakesbi /cv_for_spd_fr_2k_denoisedaudio1K<n<10K0 likes27 downloads2y agoHugging Face14kamilakesbi /cv_for_spd_ja_2k_rayleighaudio1K<n<10K0 likes24 downloads2y agoHugging Face15kamilakesbi /cv_for_spd_fr_2k_std_0.2audio1K<n<10K0 likes21 downloads2y agoHugging Face16midralab /gol-dataset-2k-ljspeechgated GOL 2K LJSpeech metadata — audited repair This gated repository contains one 320.54 GB tar archive and a pipe-delimited metadata file. The source repository did not document provenance, selection rules, audio format, license, or the meaning of “2K”. This card records only properties verified at revision 23747a89469c5487262604efb21d72bd7beef41f; it does not fill those gaps by inference. Verified contents metadata.csv: 1,652,985 logical records with contiguous IDs… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dataset-2k-ljspeech.audioautomatic-speech-recognition1M<n<10M0 likes15 downloads4d agoHugging Face17chuyangchenn /a-coat-2k A-COAT-2k Audio Compositional Object Algebra Test — 2,000 zero-shot audio quadruples for evaluating whether audio encoders represent multi-source scenes compositionally. No training required. Companion dataset to the ICASSP 2026 paper Evaluating Compositional Structure in Audio Representations. See also the trained-head benchmark chuyangchenn/a-tre-10k. Quick start from datasets import load_dataset ds = load_dataset("chuyangchenn/a-coat-2k", split="test") ex = ds[0] A… See the full description on the dataset page: https://huggingface.co/datasets/chuyangchenn/a-coat-2k.audioaudio-classification1K<n<10K0 likes15 downloads5mo agoHugging Face18xfh /ontology_image_audio_2kFrom https://github.com/audioset/ontology More Information needed audio1K<n<10K0 likes12 downloads4y agoHugging Face19nghialt /vi-songs-2k Music Query Dataset A comprehensive dataset designed to support music recognition systems. This dataset includes metadata, audio, and lyrics for 2,000 popular songs, enabling advanced music retrieval and query applications. The dataset was built by crawling and processing data from hopamchuan.com and YouTube. Dataset Contents The dataset is structured into three components: infos.jsonA JSON file containing metadata for each song. Each entry includes: Song name Author… See the full description on the dataset page: https://huggingface.co/datasets/nghialt/vi-songs-2k.audio1K<n<10K0 likes12 downloads2y agoHugging Face20amine-maazizi /tta-detection-2kaudioaudio-classification1K<n<10K0 likes12 downloads6mo agoHugging Face21mavihsrr /hindi-phoneme-2kaudio1K<n<10K0 likes11 downloads2y agoHugging Face22kamilakesbi /cv_for_spd_ja_2k_std_0.5-m0.5audio1K<n<10K0 likes10 downloads2y agoHugging Face23userdata /malaya-speech-malay-stt-2kaudio1K<n<10K0 likes8 downloads2y agoHugging Face24IbratDO /uzbekvoice-2k-each-accentaudio10K<n<100K0 likes7 downloads1y agoHugging Face25Cybrpgs /mac01-short-sbpn-low-high-2k-20260923gatedaudio1K<n<10K0 likes6 downloads3d agoHugging Face26Sam04 /yt-aud30_2k_embedded_parquetaudio100K<n<1M0 likes5 downloads10mo agoHugging Face27manishkumar2101114 /all-8-voices-2k-v2audio1K<n<10K0 likes5 downloads6mo agoHugging Face28AhunInteligence /ft_noisy_2k_normalized_rmsnorm_v1audio1K<n<10K0 likes5 downloads4mo agoHugging Face29HaofeiDing /savgbench-outputs-2kaudion<1K0 likes1 downloads7mo agoHugging Face30manishkumar2101114 /ankush-voices-2k-v2audio1K<n<10K0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.