CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CAiRE /ASCEND Dataset Card for ASCEND Dataset Summary ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/CAiRE/ASCEND.audioautomatic-speech-recognition10K<n<100K53 likes1.9k downloads2y agoHugging Face02WTFO /ascend_MIXED_cleaned_vadaudio1K<n<10K0 likes205 downloads4mo agoHugging Face03WTFO /ascend_ZH_cleaned_vadaudio10K<n<100K0 likes199 downloads4mo agoHugging Face04WTFO /ascend_EN_cleaned_vadaudio1K<n<10K0 likes193 downloads4mo agoHugging Face05georgechang8 /ASCEND_CLEAN Dataset Card for Dataset Name This dataset is derived from CAiRE/ASCEND. More information is available at https://huggingface.co/datasets/CAiRE/ASCEND. Removed 嗯 呃 um uh Resolved [UNK]'s using whisper-medium Usage Default utterances with cleaned transcripts from datasets import load_dataset data = load_dataset("georgechang8/ASCEND_CLEAN") # add split="train" for train set, etc. Concatenated 30s utterances with cleaned transcripts… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/ASCEND_CLEAN.audio10K<n<100K0 likes77 downloads2y agoHugging Face06filwsyl /ascend Dataset Card for ASCEND Dataset Summary ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/filwsyl/ascend.audioautomatic-speech-recognition1K<n<10K1 likes48 downloads4y agoHugging Face07dicksonsarpong9 /nigeria_ascentaudio1K<n<10K0 likes39 downloads9mo agoHugging Face08katyayego /ASCEND-phoneme Dataset Summary This dataset is a modified version of the ASCEND dataset which consists of spontaneous Mandarin-English code-switched speech. The ASCEND dataset was published by Lovenia et al. (2022) (Check here for the dataset and here for the paper). This dataset adds a phonetic transcription column to the dataset using the eSpeak backend from the phonemizer library created by Bernard et al. (2021) (Check it out here). the following documentation is a modified version of… See the full description on the dataset page: https://huggingface.co/datasets/katyayego/ASCEND-phoneme.audioautomatic-speech-recognition10K<n<100K1 likes32 downloads2y agoHugging Face09humanify /AS-CountingQAaudio10K<n<100K0 likes27 downloads6mo agoHugging Face10Ascyii /accent-voice-test Nyra Disfluency Speech German nyrahealth/disfluency_speech_german is a German speech dataset for evaluating verbatim ASR: models that should transcribe not only the intended words, but also fillers, cutoffs, repetitions, and sound events. This dataset was recorded in-house by two Nyra researchers, Berns and Laurin, with the goal of producing natural disfluent German speech similar in spirit to the English AMAAI Lab DisfluencySpeech dataset. Like the English release, it is… See the full description on the dataset page: https://huggingface.co/datasets/Ascyii/accent-voice-test.audioautomatic-speech-recognitionn<1K0 likes20 downloads1mo agoHugging Face11DynamicSuperb /CodeSwitchingSpeechIdentification_ASCENDaudion<1K0 likes17 downloads2y agoHugging Face12WTForbes /_ASCEND_ZH_cleanedaudio1K<n<10K0 likes17 downloads1y agoHugging Face13humanify /AS-Clotho-v2audion<1K0 likes11 downloads6mo agoHugging Face14yl31 /ASCEND-mixed-to-chinese-translationaudio1K<n<10K1 likes10 downloads2y agoHugging Face15SpeechTest /ASCENDaudio1K<n<10K0 likes10 downloads8mo agoHugging Face16WTForbes /ASCEND_ENaudio1K<n<10K0 likes8 downloads2y agoHugging Face17WTForbes /_ASCEND_MIXED_cleanedaudio1K<n<10K0 likes7 downloads1y agoHugging Face18WTForbes /ASCEND_MIXEDaudio1K<n<10K0 likes6 downloads2y agoHugging Face19WTForbes /_ASCEND_EN_cleanedaudio1K<n<10K0 likes5 downloads1y agoHugging Face20MagicLuke /ASCEND-phonemegatedaudio10K<n<100K0 likes5 downloads1y agoHugging Face21Yougen /asc_testsetgated Yougen/asc_testset Audio Scene Classification (ASC) speech dataset, packed as WebDataset tar shards. Layout data/ train/ metadata.csv audio/ train-000.tar train-001.tar ... validation/ metadata.csv audio/ validation-000.tar ... test/ metadata.csv audio/ test-000.tar ... Shard counts: test_a1: 8 tar shard(s) test_a2: 16 tar shard(s) test_a3: 13 tar shard(s) test_a4: 15 tar shard(s) test_a5: 8 tar… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/asc_testset.audioaudio-classification100K<n<1M0 likes5 downloads5mo agoHugging Face22RuishiCh0314 /ASCEND-mixed-to-chinese-translationaudio1K<n<10K1 likes3 downloads2y agoHugging Face23Yougen /asc_datasetgated Yougen/asc_dataset Audio Scene Classification (ASC) speech dataset, packed as WebDataset tar shards. Layout data/ train/ metadata.csv audio/ train-000.tar train-001.tar ... validation/ metadata.csv audio/ validation-000.tar ... test/ metadata.csv audio/ test-000.tar ... Shard counts: train: 55 tar shard(s) validation: 6 tar shard(s) Inside each tar, every sample is a pair sharing a unique key:… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/asc_dataset.audioaudio-classification10K<n<100K0 likes3 downloads5mo agoHugging Face24ellenlnt /ASCOR_audio2 Dataset Card for "ASCOR_audio2" More Information needed audion<1K0 likes2 downloads3y agoHugging Face25SST-UIUC /ASCEND-phonemegatedaudio10K<n<100K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.