datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
michii-swedish-s2-dataset
Michii Swedish Speech Dataset (Fish Speech Ready)
High-fidelity Swedish speech dataset generated via GPU faster-whisper-large-v3 with brand dictionary cleaning.
Total Audio Clips: 535
Audio Spec: 22,050 Hz, Mono, 16-bit PCM WAV
Structure: metadata.csv and .lab transcript files.
TTS-Swedish
TTS-Swedish
A high-quality Swedish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from LibriVox — Swedish audiobooks.
Dataset Statistics
Metric
Value
Total samples
14,535
Total duration
40 hours
Unique speakers
9
Average duration
10.0 seconds
Average DNSMOS
3.69
Gender Distribution
Gender
Samples
Hours
Male
11,219
31.2
Female
3,316
9.3
Features
Field… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Swedish.sweet-potatolibrivox-tts-swedish
LibriVox TTS Swedish (5-minute chunks)
Long-form variant of datadriven-company/TTS-Swedish: per-speaker audio concatenated in source order into ~5 minute chunks (16 kHz mono), for long-context TTS / ASR training.
Loading
from datasets import load_dataset
ds = load_dataset("felixmr1/librivox-tts-swedish", split="train")
# no held-out split is provided — slice it yourself, e.g. ds.train_test_split(test_size=0.05)
Columns
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/felixmr1/librivox-tts-swedish.audio_swedish_2_dataset_cleanedpop2piano_cimashup-xs-sweater-weatherSwedish-Speech-Dataset
🎧 Swedish Speech Dataset
The Swedish Speech Dataset is a high-quality speech audio dataset designed to support advanced AI and machine learning workflows with structured and diverse audio data. It contains 162 hours of voice recordings distributed across 558 files, stored in MP3 and WAV formats, with a total size of 446 MB. This carefully curated audio dataset provides rich and balanced voice data, featuring 55% female and 45% male speakers, and an age distribution ranging from 18… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Swedish-Speech-Dataset.minhavozaudio_swedish_2_datasetscotus-voice-sweep-v2Sweet-Voice-2026
Sweet Voice 2026
Access & corrections: To request access, please contact me and state your reason / intended use for this dataset. Also reach out for any correction requests.
in progress building
Single-speaker Vietnamese TTS dataset for voice cloning. 69 clips, ~4.7 minutes,
clean vocal segments (music-separated) from a single consistent voice.
Format
Audio is embedded (24 kHz mono) and plays directly in the dataset viewer.
from datasets import load_dataset
ds… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Sweet-Voice-2026.YodaLingua-Swedish
YodaLingua-Swedish
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Swedish portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
43,048 audio–transcription pairs
Total duration
112 hours
Speakers
1,946 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Swedish.
