CoolFace
Datasetpublic

AudioCC-Lab/PICSAFEv1

Speech Quality Test Labels It publishes only this README and metadata.jsonl, because the 14 source datasets have different copyright and access terms. The manifest contains the 10,728-sample PICSAFEv1 subset. It is not a license to redistribute the underlying recordings. Metadata fields file_name: audio basename. id: globally unique sample identifier with the [source] prefix. source: source dataset name. source_path: path relative to the origin directory. tags:… See the full description on the dataset page: https://huggingface.co/datasets/AudioCC-Lab/PICSAFEv1.

sourceHugging Faceotherupdated 4d agoView on Hugging Face
0likes67downloads
Dataset Card

Speech Quality Test Labels

It publishes only this README and `metadata.jsonl`, because the 14 source datasets have different copyright and access terms.

The manifest contains the 10,728-sample PICSAFEv1 subset. It is not a license to redistribute the underlying recordings.

Metadata fields

  • file_name: audio basename.
  • id: globally unique sample identifier with the [source] prefix.
  • source: source dataset name.
  • source_path: path relative to the origin directory.
  • tags: list of labels.

Source datasets and official download paths

Source datasetSamplesLicense / access noteOfficial download or access path
CSTR-NAM-TIMIT-Plus842ODC-By 1.0Edinburgh DataShare
DNS5_LibriVox510LibriVox public domain; jurisdiction restrictions may applyMicrosoft DNS5 download instructions
EARS1,042CC BY-NC 4.0; non-commercial useOfficial GitHub repository
MSceneSpeech1,000Dataset audio terms are not stated separately; confirm with authorsProject page / Google Drive
NISQA (NISQATESTFOR)724Original source terms; this subset is restricted to non-commercial research/forensic useNISQA Corpus wiki / DepositOnce archive
SOMOSv2500Research and non-commercial use only; redistribution under the same termsZenodo record
TencentCorpus1,010No separate public audio license located; challenge access terms applyConferencingSpeech2022 repository / challenge plan
VCTK1,000CC BY 4.0 for the official VCTK release; verify the exact versionEdinburgh DataShare
WenetSpeech600CC BY 4.0; access requires the official form/password workflowOpenSLR SLR121
Whisper40500No explicit dataset license located in the official repositoryOfficial GitHub repository
WSJ1,000LDC User Agreement; license requiredWSJ0 / LDC93S6A and WSJ1 / LDC94S13A
zhvoice500Mixed upstream datasets; no unified licenseOfficial GitHub repository (Baidu download link is on the page)
LibriTTS-R500CC BY 4.0OpenSLR SLR141
LibriTTS1,000CC BY 4.0OpenSLR SLR60

Label schema

In article, it is described that the 33 labels are:

Nspk, distortion, emotional, enhanced, excessive_sibilance, fast_speed, female, high_pitch, imperceptible_noise_level, instantaneous_noise, low_noise_level, low_pitch, male, microphone_popping, mid_high_noise_level, music_or_effect, nb_noise, non_binary, non_speech, read_speech, real_recording, regular_pitch, regular_speed, reverberation, singing_voice, slow_speed, speech_overlap, spontaneous_speech, synthetic, vocal_sound, wb_noise, whispered_speech, and wrong_transcript.

Licensing and limitations

This repository does not grant, aggregate, or override any source-dataset license. Obtain the audio directly from the source providers and follow their current terms, attribution requirements, access restrictions, and applicable privacy/publicity rules. Some sources require non-commercial research use, registration, a license, or direct permission from the authors.