datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TTS-Greek
TTS-Greek
A large-scale, high-quality Greek speech dataset for text-to-speech and automatic speech recognition.
Data Sources
This dataset combines two sources:
Source
Samples
Hours
License
Content
LibriVox
34,727
96.8
Public Domain
Modern Greek classic literature, philosophy, fiction
FLEURS-R (Google)
4,124
12.6
CC-BY 4.0
Wikipedia-sourced sentences, AI-restored audio
Dataset Statistics
Metric
Value
Total samples
38,851
Total… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Greek.Greek-Speech-Dataset
🎧 Greek Speech Dataset
The Greek Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that rely on diverse audio data and multilingual voice data. It contains 184 hours of recordings distributed across 592 files, stored in MP3 and WAV formats, with a total size of 330 MB. This carefully curated audio dataset delivers balanced representation across speakers, including 49% female and 51% male participants, and a broad age range… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Greek-Speech-Dataset.YodaLingua-Greek
YodaLingua-Greek
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Greek portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
28,839 audio–transcription pairs
Total duration
81 hours
Speakers
1,658 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Greek.
