datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-voicesamples-generatedHy-Generated-audio-data-with-cv20.0
Hy-Generated Audio Data with CV20.0
This dataset provides Armenian speech data consisting of both real and generated audio clips.
The train, test, and eval splits are derived from the Common Voice 20.0 Armenian dataset.
The generated split contains 100,000 high-quality clips synthesized using a fine-tuned F5-TTS model, covering 404 equal distribution of synthetic voices.
📊 Dataset Statistics
Split
# Clips
Duration (hours)
train
9,300
13.53
test
5,818… See the full description on the dataset page: https://huggingface.co/datasets/ErikMkrtchyan/Hy-Generated-audio-data-with-cv20.0.Hy-Generated-audio-data-2
Hy-Generated Audio Data 2
This dataset provides Armenian speech data consisting of generated audio clips and is addition to this dataset.
The generated split contains 137,419 high-quality clips synthesized using a fine-tuned F5-TTS model, covering 404 equal distribution of synthetic voices.
📊 Dataset Statistics
Split
# Clips
Duration (hours)
generated
137,419
173.76
Total duration: ~173 hours
🛠️ Loading the Dataset
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ErikMkrtchyan/Hy-Generated-audio-data-2.ATC_generated_realATC_TTS_generatedvoxcpm2-native-generated-audio-user-ref
VoxCPM2 Native Generated Audio (User Ref)
Raw audio files generated from the native VoxCPM2 path in sglang-omni using a user-provided reference clip.
Contents
9 generated .wav files
metadata.json with prompt text, mode, status, size, and latency
Source Reference Audio
Reference clip used for the reference-mode generations:
https://huggingface.co/datasets/adarshxs/voxcpm2-native-test-samples/resolve/main/data/audio.wav
Files
ref_expressive.wav… See the full description on the dataset page: https://huggingface.co/datasets/adarshxs/voxcpm2-native-generated-audio-user-ref.datasets-asr-generateai-generated-songsdiffrhythm-instrument-cc0-oepngamearg-10x5-generated
What is this
A dataset of 50 instrumental music tracks generated with the DiffRhythm model, using 10 CC0-licensed instrument samples from OEPN Game Art.
Paper: https://huggingface.co/papers/2503.01183
Project page: https://nzqian.github.io/DiffRhythm/
Models
This dataset was created by running the DiffRhythm model on 2025 Mar 05 using a copy of the space available at https://huggingface.co/spaces/ASLP-lab/DiffRhythm.
The exact architecture (base or VAE) of the DiffRhythm… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/diffrhythm-instrument-cc0-oepngamearg-10x5-generated.ai-generated-podcast-episodesmusic_generate_baselineai-generated-songs2banger-scorer-generated-songs
Banger Scorer Generated Songs
230 AI-generated songs across 10 genres and 5 languages, all scored by the banger scorer. Includes 10 genre tests (20 songs each) plus 1 banger-optimized run (30 songs) that used data-driven parameter selection to maximize scores.
Every song includes its MP3 audio, generation metadata (BPM, key, seed, caption prompt), and banger score. Useful for research into AI music quality, training better scorers, or just listening to what works and what does not.… See the full description on the dataset page: https://huggingface.co/datasets/treadon/banger-scorer-generated-songs.ai-generated-songs3khmer-ASR-Generate-AIphone-recognition-generated
Dataset Card for "phone-recognition-generated"
More Information needed
generated-sound-eventsSLIDE_generatedtts-generated-audio-orpheusGov-ASR-Generateazure_generate_thai_audio_10000
Dataset Card for "azure_generate_thai_audio_10000"
More Information needed
GeneratedMusicazure_generate_thai_audio_800
Dataset Card for "azure_generate_thai_audio_800"
More Information needed
Test_Audio_Generate_Dataset
Hinglish Audio Dataset
Generated by Sarvam AI.
tts-generated-audioAudioX_generated_audiogenerated-audio-vocoder-MelGAN-HiFiGAN-PWG-MultibandMelGAN-FullbandMelGAN-WaveGlowThis dataset includes samples of the vocoder of MelGAN, HiFi-GAN, Parallel WaveGAN (PWG), Multi-band MelGAN, Full-band MelGAN, and WaveGlow.
datasets-asr-generategenerated-26k-audiogenerated
