datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASCEND
Dataset Card for ASCEND
Dataset Summary
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/CAiRE/ASCEND.ascend_MIXED_cleaned_vadascend_ZH_cleaned_vadascend_EN_cleaned_vadASCEND_CLEAN
Dataset Card for Dataset Name
This dataset is derived from CAiRE/ASCEND. More information is available at https://huggingface.co/datasets/CAiRE/ASCEND.
Removed 嗯 呃 um uh
Resolved [UNK]'s using whisper-medium
Usage
Default utterances with cleaned transcripts
from datasets import load_dataset
data = load_dataset("georgechang8/ASCEND_CLEAN") # add split="train" for train set, etc.
Concatenated 30s utterances with cleaned transcripts… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/ASCEND_CLEAN.ascend
Dataset Card for ASCEND
Dataset Summary
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/filwsyl/ascend.nigeria_ascentASCEND-phoneme
Dataset Summary
This dataset is a modified version of the ASCEND dataset which consists of spontaneous Mandarin-English code-switched speech. The ASCEND dataset was published by Lovenia et al. (2022) (Check here for the dataset and here for the paper).
This dataset adds a phonetic transcription column to the dataset using the eSpeak backend from the phonemizer library created by Bernard et al. (2021) (Check it out here).
the following documentation is a modified version of… See the full description on the dataset page: https://huggingface.co/datasets/katyayego/ASCEND-phoneme.AS-CountingQAaccent-voice-test
Nyra Disfluency Speech German
nyrahealth/disfluency_speech_german is a German speech dataset for evaluating verbatim ASR: models that should transcribe not only the intended words, but also fillers, cutoffs, repetitions, and sound events.
This dataset was recorded in-house by two Nyra researchers, Berns and Laurin, with the goal of producing natural disfluent German speech similar in spirit to the English AMAAI Lab DisfluencySpeech dataset.
Like the English release, it is… See the full description on the dataset page: https://huggingface.co/datasets/Ascyii/accent-voice-test.CodeSwitchingSpeechIdentification_ASCEND_ASCEND_ZH_cleanedAS-Clotho-v2ASCEND-mixed-to-chinese-translationASCENDASCEND_EN_ASCEND_MIXED_cleanedASCEND_MIXED_ASCEND_EN_cleanedASCEND-phonemeasc_testset
Yougen/asc_testset
Audio Scene Classification (ASC) speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
test_a1: 8 tar shard(s)
test_a2: 16 tar shard(s)
test_a3: 13 tar shard(s)
test_a4: 15 tar shard(s)
test_a5: 8 tar… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/asc_testset.ASCEND-mixed-to-chinese-translationasc_dataset
Yougen/asc_dataset
Audio Scene Classification (ASC) speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
train: 55 tar shard(s)
validation: 6 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/asc_dataset.ASCOR_audio2
Dataset Card for "ASCOR_audio2"
More Information needed
ASCEND-phoneme
