CoolFace
Datasetpublic

icfoss/malayalam-asr-5K

Malayalam ASR 5K — Verified Anchor Set 5,225 manually verified Malayalam speech-transcript pairs (7.1 hours), speaker-disjoint across train/dev/test. Every record in this release carries is_verified: true — each transcript was checked, not machine- generated and left unreviewed. Splits (speaker-disjoint, source-aware) split utterances % disjoint units speakers hours train 3,657 70% 9 7 4.8 dev 783 15% 5 5 1.1 test 785 15% 6 5 1.1 No speaker or… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-asr-5K.

sourceHugging Faceunknownupdated 10d agoView on Hugging Face
0likes94downloads
Dataset Card

Malayalam ASR 5K — Verified Anchor Set

5,225 manually verified Malayalam speech-transcript pairs (7.1 hours), speaker-disjoint across train/dev/test. Every record in this release carries is_verified: true — each transcript was checked, not machine- generated and left unreviewed.

Splits (speaker-disjoint, source-aware)

splitutterances%disjoint unitsspeakershours
train3,65770%974.8
dev78315%551.1
test78515%651.1

No speaker or source group appears in more than one split, and there is zero identical-transcript overlap across splits (train∩dev = train∩test = dev∩test = 0).

Fields

Each split folder (train/, dev/, test/) contains the audio files plus a metadata.csv with columns:

  • —file_name — audio filename (mono WAV, 44.1kHz)
  • —transcript — verified Malayalam transcription
  • —speaker_id — anonymized speaker identifier (e.g. SPEAKER_08)
  • —source_dataset — originating recording/source batch
  • —duration — clip duration in seconds (range: 2.0–10.0s)

Quality control

  • —18 distinct speakers, 15 distinct sources.
  • —5,247 records considered before deduplication; 5,225 unique transcripts remain after removing 15 duplicate transcripts (37 flagged occurrences).
  • —All audio confirmed present and readable at a consistent 44.1kHz sample rate, mono channel.

License

Not explicitly specified by the source data; released here by ICFOSS. Please contact ICFOSS regarding reuse terms if you have questions.

Acknowledgments

With gratitude to everyone at ICFOSS who worked on building the Malayalam speech corpus this anchor set is drawn from — the recording, annotation, and verification effort that makes a genuinely human-checked dataset like this possible.

Citation

If you use this dataset, please credit ICFOSS (Indian Institute of Free and Open Source Software).