datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv-corpus-17.0-zh-TW-client_id-grouped
cv-corpus-17.0-zh-TW-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-TW-client_id-grouped.cv-corpus-1.0-en-client_id-grouped
cv-corpus-1.0-en-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.cv-corpus-17.0-zh-CN-client_id-grouped
cv-corpus-17.0-zh-CN-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-CN-client_id-grouped.cv-corpus-17.0-ja-client_id-grouped
cv-corpus-17.0-ja-client_id-grouped
This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID).
Dataset Details
The dataset is derived from the Common Voice dataset.
The original dataset is available at Common Voice Dataset.
The dataset is grouped by client ID, which is treated as the speaker ID for this dataset.
Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-ja-client_id-grouped.rapnic-example
RAPNIC Dataset (example)
Dataset Description
This is an example of the full dataset, yet to be published, with 10 audio examples for 72 speakers.
RAPNIC (Reconeixement Automàtic de la Parla No Intel·ligible en Català) is a Catalan speech corpus collected from individuals with speech disorders, specifically cerebral palsy and Down syndrome.
This dataset was collected to develop and improve automatic speech recognition (ASR) systems that are accessible to people with speech… See the full description on the dataset page: https://huggingface.co/datasets/CLiC-UB/rapnic-example.
