datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DNS-Noise
DNS-Noise
DNS-Challenge noise_fullband and impulse_responses republished as 24 kHz mono
FLAC for on-demand streaming augmentation, plus a microphone-characteristics
split. Three splits:
noise — environmental noise (long files segmented into chunks)
rir — room impulse responses (one row per file)
mic_ir — microphone impulse responses (one row per file)
from datasets import load_dataset
noise = load_dataset("ChristianYang/DNS-Noise", split="noise", streaming=True)
rir =… See the full description on the dataset page: https://huggingface.co/datasets/humanify/DNS-Noise.dns5
DNS5 Challenge data
This is a mirror of the DNS5 Challenge data.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
⚠️ Only the LibriVox, AudioSet, Freesound, OpenSLR26, and OpenSLR28 data is included. The VCTK, VocalSet, CREMA-D, VoxCeleb2, and DEMAND data is excluded. ⚠️
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
speech_english: 245 hours, 186743 files
speech_french: 95 hours, 60454 files
speech_german: 137 hours, 119175… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/dns5.dns-recordsdns5-16k
DNS5 16kHz
Resampled subset of the ICASSP 2022 DNS Challenge dataset.
All audio files resampled from 48kHz to 16kHz and stored as FLAC (lossless compression),
packed into tar shards.
Structure
clean/shard_0000.tar # Clean speech (VCTK and other corpora)
clean/shard_0001.tar
...
noise/shard_0000.tar # Environmental noise (AudioSet, Freesound)
...
impulse_responses/shard_0000.tar # Room impulse responses
...
Each tar contains FLAC files with their… See the full description on the dataset page: https://huggingface.co/datasets/richiejp/dns5-16k.DNS-Challenge-2020-DevTest-16k
DNS Challenge 2020 Dev Test Set
Preprocessed dev test set from the Interspeech 2020 DNS Challenge.
Dataset Structure
Split
Samples
Clean Reference
Description
synthetic_no_reverb
150
✓
Anechoic synthetic mixtures
synthetic_with_reverb
150
✓
Reverberant synthetic mixtures
real_recordings
300
✗
Real-world noisy recordings
Usage
from datasets import load_dataset
ds = load_dataset("nkdem/DNS-Challenge-2020-DevTest-16k")
# Access a sample… See the full description on the dataset page: https://huggingface.co/datasets/nkdem/DNS-Challenge-2020-DevTest-16k.denoised-reazonspeech-v2-dnsmosDNSMOS score of reazon-speech-v2-denoised
dn_sft_part_7dn_sft_part_5dn_sft_part_2dn_sft_part_1Multilingual-TTS-DNSMOS
Multilingual TTS — DNSMOS Filtered
A quality-filtered subset of malaysia-ai/Multilingual-TTS, retaining only audio samples that score OVRL ≥ 3.2 on the DNSMOS non-intrusive speech quality metric.
4,911,906 samples across multiple languages and TTS sources (84++ subset) remain after filtering.
Dataset Structure
Each record in the dataset contains the following fields:
Column
Type
Description
audio_filename
string
Relative path to the audio file within its… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-TTS-DNSMOS.dn_sft_part_4RTSP_FTP_SMTP_DNS_PROTOCOL_GRAMMAR_DATASETDoD-Instruction-8410-DNS-IP-Address-Use-And-Approval
🌐 DoD Internet Domain Name and IP Address Resource Question-Answer Dataset
Source: DoD Instruction 8410.01
Source Effective Date: December 4, 2015
Change Incorporated: Change 1, effective June 4, 2021
Source Organization: Office of the DoD Chief Information Officer
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Internet Domain Name and IP Address Resource Question-Answer Dataset is a structured… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8410-DNS-IP-Address-Use-And-Approval.network-dns-resolution-coherence-risk-v0.1What this repo is for
Detect DNS instability before services fail.
Covers real operational signals:
rising resolution latency
SERVFAIL spikes
authoritative mismatch
cache poisoning
missing failover resolvers
Used by:
ISPs
cloud providers
enterprises
SRE teams
dnsdn_sft_part_0DNSMOS-TTS(Placeholder)
DNSMOS-TTS
DNSMOS-TTS contains DNSMOS Scores for common TTS datasets
This repo uses Lhotse to manage datasets.
For example, to load LJ-Speech:
from lhotse import CutSet
for cut in CutSet.from_webdataset("pipe:curl -s -L https://huggingface.co/datasets/Gatozu35/DNSMOS-TTS/resolve/main/ljspeech_mos.tar"):
wav = cut.load_audio()
mos = cut.supervisions[0].custom["mos"]
...
If you don't want to use lhotse, I have also uploaded a csv of the scores for each id.
dnsy-1-qposneg-3testing-dnsmosserial2023
Dataset Card for "serial2023"
More Information needed
dnsy-1-qposneg-2dn_sft_part_3dn_sft_part_6dnsy-2serials
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Dnsibu/serials.dnsy-1sn
Dataset Card for "sn"
More Information needed
voice-design-bench-50-dnsmos
Model
Mean
Min
Max
Qwen
3.262
2.480
3.619
Echo
3.217
2.158
3.588
Omni
3.190
0.988
3.629
sn_tsv
