datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icelandic_asr
Icelandic ASR Collection
This repository collects six Icelandic speech corpora in directly loadable
Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a
convenience repackaging: the linked CLARIN-IS records and original dataset
repositories remain the canonical sources and should be cited when using the
data.
No configuration is selected by default. Choose a corpus configuration and,
for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.stortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.
