icfoss/malayalam-asr-5K
Malayalam ASR 5K — Verified Anchor Set 5,225 manually verified Malayalam speech-transcript pairs (7.1 hours), speaker-disjoint across train/dev/test. Every record in this release carries is_verified: true — each transcript was checked, not machine- generated and left unreviewed. Splits (speaker-disjoint, source-aware) split utterances % disjoint units speakers hours train 3,657 70% 9 7 4.8 dev 783 15% 5 5 1.1 test 785 15% 6 5 1.1 No speaker or… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-asr-5K.
Malayalam ASR 5K — Verified Anchor Set
5,225 manually verified Malayalam speech-transcript pairs (7.1 hours), speaker-disjoint across train/dev/test. Every record in this release carries is_verified: true — each transcript was checked, not machine- generated and left unreviewed.
Splits (speaker-disjoint, source-aware)
No speaker or source group appears in more than one split, and there is zero identical-transcript overlap across splits (train∩dev = train∩test = dev∩test = 0).
Fields
Each split folder (train/, dev/, test/) contains the audio files plus a metadata.csv with columns:
file_name— audio filename (mono WAV, 44.1kHz)transcript— verified Malayalam transcriptionspeaker_id— anonymized speaker identifier (e.g.SPEAKER_08)source_dataset— originating recording/source batchduration— clip duration in seconds (range: 2.0–10.0s)
Quality control
- 18 distinct speakers, 15 distinct sources.
- 5,247 records considered before deduplication; 5,225 unique transcripts remain after removing 15 duplicate transcripts (37 flagged occurrences).
- All audio confirmed present and readable at a consistent 44.1kHz sample rate, mono channel.
License
Not explicitly specified by the source data; released here by ICFOSS. Please contact ICFOSS regarding reuse terms if you have questions.
Acknowledgments
With gratitude to everyone at ICFOSS who worked on building the Malayalam speech corpus this anchor set is drawn from — the recording, annotation, and verification effort that makes a genuinely human-checked dataset like this possible.
Citation
If you use this dataset, please credit ICFOSS (Indian Institute of Free and Open Source Software).
