CoolFace
Datasetpublic

Speech-data/Greek-Speech-Dataset

๐ŸŽง Greek Speech Dataset The Greek Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that rely on diverse audio data and multilingual voice data. It contains 184 hours of recordings distributed across 592 files, stored in MP3 and WAV formats, with a total size of 330 MB. This carefully curated audio dataset delivers balanced representation across speakers, including 49% female and 51% male participants, and a broad ageโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Greek-Speech-Dataset.

sourceHugging Facecc-by-nc-nd-4.0updated 6mo agoView on Hugging Face
0likes22downloads
Dataset Card

๐ŸŽง Greek Speech Dataset

The Greek Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that rely on diverse audio data and multilingual voice data. It contains 184 hours of recordings distributed across 592 files, stored in MP3 and WAV formats, with a total size of 330 MB. This carefully curated audio dataset delivers balanced representation across speakers, including 49% female and 51% male participants, and a broad age range from 18 to 50+ years. The dataset language is Greek, with contributions from speakers across Greece, Cyprus, the USA, Australia, Germany, and the UK, ensuring a globally representative language speech dataset.


๐Ÿ”— Learn more: https://speech-data.ai/datasets/greek/


๐Ÿš€ Use Cases

This Greek speech dataset is designed for advanced AI and NLP applications, including speech recognition, voice assistant development, and natural language processing. The structured speech data supports acoustic modeling, speaker identification, and robust AI training pipelines. It also enhances multilingual system development where reliable speech audio dataset resources are required. As a high-quality speech recognition dataset, it is suitable for both academic research and production-level deployment in voice technologies.


๐Ÿ“Š Dataset Metadata

FieldValue
๐Ÿ“œ LicenseCC BY-NC-ND 4.0
๐ŸŽฏ Task CategoriesAutomatic Speech Recognition
๐ŸŒ LanguageGreek (el)
๐Ÿท๏ธ TagsAudio, Speech, Speech Recognition, Greek, Machine, Machine Learning
๐Ÿ“ฆ Size Categoryn < 1K

โญ Key Value

The key strength of this voice dataset lies in its linguistic diversity, balanced demographic structure, and high-quality audio data suitable for real-world AI training. It improves model robustness across accents and regional variations, making it a valuable resource for scalable speech dataset development and multilingual speech technology systems.