datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.callcc-2k-hours-train-soniox
ErfanRou/callcc-2k
Persian call-centre ASR silver set (2,000 audio-hours, 2,288 h shipped): the ErfanRou/callcc-ft150-soniox schema plus two auxiliary columns.
Machine-labelled, not ground truth. Built by callcc-silver-5k/soniox_pipeline.py from markmuller/call-center-prod-data:
8 kHz stereo telephony split client-side into agent = channel 0 and customer = channel 1, each channel transcribed
separately by Soniox stt-async-v5, words packed into 5-28 s windows (15 s mean, natural… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/callcc-2k-hours-train-soniox.cv_for_spd_fr_2k_augmentedcv_for_spd_fr_augmented_2kndizi-1_sample_2kkaito_2kall-12-voices-2kradiotalk-voices-2k
radiotalk-voices-2k
2,000 English reference voices — one 12–30s clip per speaker, selected as the longest qualifying utterance per speaker from LibriTTS-R. Built for zero-shot TTS voice cloning in the radiotalk pipeline.
Stats
2,000 voices · 12.03 hours total
Duration: min 12.0s · median 21.8s · mean 21.7s · max 30.0s
24 kHz, mono, FLAC-encoded
Schema
Column
Type
Description
voice_id
string
Stable 12-hex-char id, derived from (source… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-voices-2k.cv_for_spd_fr_2k_std_0.5text-vision-audio-2k-testA 2k sample dataset for testing multimodal (text+vision+audio) format. This is compatible with HF's processor apply_chat_template.
Load in Axolotl via:
datasets:
- path: Nanobit/text-vision-audio-2k-test
type: chat_template
Make sure to download the image and audio via:
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/African_elephant.jpg
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/En-us-African_elephant.oga… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-audio-2k-test.text-audio-2k-testA 2k sample dataset for testing multimodal (text+audio) format. This is compatible with HF's processor apply_chat_template.
Load in Axolotl via:
datasets:
- path: Nanobit/text-audio-2k-test
type: chat_template
Make sure to download the audio via:
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/En-us-African_elephant.oga
Audio source: https://upload.wikimedia.org/wikipedia/commons/a/ad/En-us-African_elephant.oga
Each sample has the following format… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-audio-2k-test.Hindi_TTS_M-2kcv_for_spd_fr_2k_denoisedcv_for_spd_ja_2k_rayleighcv_for_spd_fr_2k_std_0.2gol-dataset-2k-ljspeech
GOL 2K LJSpeech metadata — audited repair
This gated repository contains one 320.54 GB tar archive and a pipe-delimited metadata
file. The source repository did not document provenance, selection rules, audio format,
license, or the meaning of “2K”. This card records only properties verified at revision
23747a89469c5487262604efb21d72bd7beef41f; it does not fill those gaps by inference.
Verified contents
metadata.csv: 1,652,985 logical records with contiguous IDs… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dataset-2k-ljspeech.a-coat-2k
A-COAT-2k
Audio Compositional Object Algebra Test — 2,000 zero-shot audio
quadruples for evaluating whether audio encoders represent multi-source scenes
compositionally. No training required.
Companion dataset to the ICASSP 2026 paper Evaluating Compositional Structure in Audio
Representations. See also the
trained-head benchmark chuyangchenn/a-tre-10k.
Quick start
from datasets import load_dataset
ds = load_dataset("chuyangchenn/a-coat-2k", split="test")
ex = ds[0]
A… See the full description on the dataset page: https://huggingface.co/datasets/chuyangchenn/a-coat-2k.ontology_image_audio_2kFrom https://github.com/audioset/ontology
More Information needed
vi-songs-2k
Music Query Dataset
A comprehensive dataset designed to support music recognition systems. This dataset includes metadata, audio, and lyrics for 2,000 popular songs, enabling advanced music retrieval and query applications. The dataset was built by crawling and processing data from hopamchuan.com and YouTube.
Dataset Contents
The dataset is structured into three components:
infos.jsonA JSON file containing metadata for each song. Each entry includes:
Song name
Author… See the full description on the dataset page: https://huggingface.co/datasets/nghialt/vi-songs-2k.tta-detection-2khindi-phoneme-2kcv_for_spd_ja_2k_std_0.5-m0.5malaya-speech-malay-stt-2kuzbekvoice-2k-each-accentmac01-short-sbpn-low-high-2k-20260923yt-aud30_2k_embedded_parquetall-8-voices-2k-v2ft_noisy_2k_normalized_rmsnorm_v1savgbench-outputs-2kankush-voices-2k-v2
