datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_20_armenian
Common Voice 20 - Armenian
This dataset is the Armenian portion of Mozilla's Common Voice 20.0 release,
a massively multilingual collection of transcribed speech intended for speech technology research and development.
Dataset Details
Language: Armenian (hy)
Source: Mozilla Common Voice
Version: 20.0
License: CC0-1.0
armenian-speech-dataset
🎧 Armenian Speech Dataset
📘 Overview
The Armenian Speech Dataset is a high-quality speech audio dataset designed for building, training, and evaluating modern AI voice technologies. It provides structured audio data optimized for deep learning workflows in speech processing. The dataset includes 76 hours of audio data distributed across 558 files, delivered in MP3 and WAV formats, with a total size of 189 MB.
This carefully curated audio dataset ensures balanced and… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/armenian-speech-dataset.
