CoolFace
Datasetpublic

psk/malayalam-speech-178h

Malayalam Speech — 179 hours (denoised, unlabeled) 86,799 Malayalam speech clips, 48 kHz stereo WAV, ~7.4 s average. No transcripts — this is unlabeled audio, intended for self-supervised pretraining, voice/speaker modelling, or as raw material for your own labelling pipeline. Processing Each clip passed through a full source-separation and enhancement chain: Vocal isolation — BS-RoFormer De-reverberation — UVR-DeEcho-DeReverb Noise removal — UVR-DeNoise-Lite… See the full description on the dataset page: https://huggingface.co/datasets/psk/malayalam-speech-178h.

sourceHugging Faceupdated 2mo agoView on Hugging Face
4likes422downloads
Dataset Card

Malayalam Speech — 179 hours (denoised, unlabeled)

86,799 Malayalam speech clips, 48 kHz stereo WAV, ~7.4 s average. No transcripts — this is unlabeled audio, intended for self-supervised pretraining, voice/speaker modelling, or as raw material for your own labelling pipeline.

Processing

Each clip passed through a full source-separation and enhancement chain:

  1. 1.Vocal isolation — BS-RoFormer
  2. 2.De-reverberation — UVR-DeEcho-DeReverb
  3. 3.Noise removal — UVR-DeNoise-Lite
  4. 4.Final enhancement

Clips shorter than 1 second were dropped. Clip IDs are randomly generated and carry no ordering or grouping information.

Source and licensing

The audio is derived from publicly available Malayalam video content on the internet, segmented and processed as described above. Original speakers did not consent to inclusion and the underlying recordings may be subject to copyright. It is published here for research use; verify your own legal position before using it commercially or redistributing it. If you hold rights to material in this dataset and want it removed, open a discussion on this repository.

Columns

  • audio — 48 kHz stereo waveform
  • clip_id — randomized identifier
  • duration — clip length in seconds