datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.STT_uzThe dataset is organized into the following directories and files:
audio/
other/: Contains .tar archives like uz_other_0.taruz_other_1.tar
train/: Contains .tar archives like uz_train_0.tar.
validated/: Contains .tar archives like uz_validated_0.tar, uz_validated_1.tar, and uz_validated_2.tar.
test/: Contains individual .wav files.
transcription/: Contains .tsv files including:
other.tsv
train.tsv
validated.tsv
test.tsv
The .tsv files have two columns: file_name and transcription. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/Beehzod/STT_uz.uzbek_stt_data
