Center-of-Advanced-Software-Technologies/mwa_hy
Modern Western Armenian dataset An audio–transcription alignment dataset for the Modern Western Armenian dialect of Armenian. It contains approximately 42 hours of speech distributed across three splits: Train: 13,271 samples / 41 hr. Validation (Dev): 120 samples / 0.3 hr. Test: 178 samples / 0.5 hr. Content Each example includes: audio: a WAV audio file transcription: the Armenian transcription text duration: audio duration in seconds Sources The recordings were sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/Center-of-Advanced-Software-Technologies/mwa_hy.
Modern Western Armenian dataset
An audio–transcription alignment dataset for the Modern Western Armenian dialect of Armenian.
It contains approximately 42 hours of speech distributed across three splits:
- Train: 13,271 samples / 41 hr.
- Validation (Dev): 120 samples / 0.3 hr.
- Test: 178 samples / 0.5 hr.
Content
Each example includes:
- audio: a WAV audio file
- transcription: the Armenian transcription text
- duration: audio duration in seconds
Sources
The recordings were sourced from the ReRooted-ArmenianCorpus, a manually annotated Modern Western Armenian speech corpus, available on GitHub: https://github.com/jhdeov/ReRooted-ArmenianCorpus. Also from the website https://www.rerooted.org/
