CoolFace
Datasetpublic

Center-of-Advanced-Software-Technologies/mwa_hy

Modern Western Armenian dataset An audio–transcription alignment dataset for the Modern Western Armenian dialect of Armenian. It contains approximately 42 hours of speech distributed across three splits: Train: 13,271 samples / 41 hr. Validation (Dev): 120 samples / 0.3 hr. Test: 178 samples / 0.5 hr. Content Each example includes: audio: a WAV audio file transcription: the Armenian transcription text duration: audio duration in seconds Sources The recordings were sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/Center-of-Advanced-Software-Technologies/mwa_hy.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
0likes7downloads
Dataset Card

Modern Western Armenian dataset

An audio–transcription alignment dataset for the Modern Western Armenian dialect of Armenian.

It contains approximately 42 hours of speech distributed across three splits:

  • —Train: 13,271 samples / 41 hr.
  • —Validation (Dev): 120 samples / 0.3 hr.
  • —Test: 178 samples / 0.5 hr.

Content

Each example includes:

  • —audio: a WAV audio file
  • —transcription: the Armenian transcription text
  • —duration: audio duration in seconds

Sources

The recordings were sourced from the ReRooted-ArmenianCorpus, a manually annotated Modern Western Armenian speech corpus, available on GitHub: https://github.com/jhdeov/ReRooted-ArmenianCorpus. Also from the website https://www.rerooted.org/