audio-quality
subjective_audio_quality
Balanced Perceptual Audio Quality Dataset
Dataset Summary
This is a large-scale, balanced dataset designed for training models for perceptual audio quality assessment. It consists of 612,020 examples, each containing a pair of 1-second audio clips: a high-quality original and a degraded version processed by various audio codecs. Each pair is accompanied by a perceptual quality score (ranging from 0.0 to 1.0) generated by visqol-like algorithms.
The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/overfitprolabse/subjective_audio_quality.audio-quality-dataset-nfe4-30-step2
Audio Quality Dataset: NFE 4-30 Step 2
Overview
This dataset publishes synthetic speech artifacts and derived spectrograms used for repo-local audio-quality experiments.
At a glance:
2800 synthetic runs
200 short English prompt sentences
14 NFE settings: 4, 6, 8, ..., 30
fixed seed 1024
Each row represents one synthetic run and includes:
prompt text
raw synthetic WAV
processed synthetic WAV
spectrogram PNG
NFE value
procedural weak label
Here, NFE means the number of… See the full description on the dataset page: https://huggingface.co/datasets/TashaSkyUp/audio-quality-dataset-nfe4-30-step2.common_voice_audio_quality_enhancement_v3audio-quality-whitepaper
Audio Quality Whitepaper
DNSMOS is the industry standard for scoring speech audio quality. It's also not built to tell you if your training data is right for your specific model.
Most real-world audio doesn't cleanly pass or fail. It lands in the grey zone between a DNSMOS of 3 and 4. Cut it, and you're discarding usable data. Keep it without calibration, and you risk quality regressions.
This whitepaper covers the calibration framework we use, how we validate audio quality… See the full description on the dataset page: https://huggingface.co/datasets/voicesinc/audio-quality-whitepaper.common_voice_audio_quality_enhancementIndicVoices_Hindi_audio_44100_18_30_other_quality_metadata
