datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malayalam_CultureX_IndicCorp_SMCMalayalam Pretraining/Tokenization dataset.
Preprocessed and combined data from the following links,
* ai4bharat
* CulturaX
* Swathanthra Malayalam Computing
Commands used for preprocessing.
To remove all non Malayalam characters.
sed -i 's/[^ം-ൿ.,;:@$%+&?!() ]//g' test.txt
To merge all the text files in a particular Directory(Sub-Directory)
find SMC -type f -name '*.txt' -exec cat {} ; >> combined_SMC.txt
To remove all lines with characters less than 5.
grep -P… See the full description on the dataset page: https://huggingface.co/datasets/VishnuPJ/Malayalam_CultureX_IndicCorp_SMC.asr_malayalamOCR-bench-Malayalammalayalam_2020_wiki��This dataset is from the common-crawl-malayalam repo: https://github.com/qburst/common-crawl-malayalam
indic-Malayalam-PDSPRING_INX_Malayalam_R1malayalam-speech-178h
Malayalam Speech — 179 hours (denoised, unlabeled)
86,799 Malayalam speech clips, 48 kHz stereo WAV, ~7.4 s average.
No transcripts — this is unlabeled audio, intended for self-supervised pretraining,
voice/speaker modelling, or as raw material for your own labelling pipeline.
Processing
Each clip passed through a full source-separation and enhancement chain:
Vocal isolation — BS-RoFormer
De-reverberation — UVR-DeEcho-DeReverb
Noise removal — UVR-DeNoise-Lite… See the full description on the dataset page: https://huggingface.co/datasets/psk/malayalam-speech-178h.malayalam-ocr-words
Malayalam OCR Words
A word-level Malayalam OCR dataset: cropped word images paired with their
transcribed text label, split into train/validation/test sets.
Dataset structure
train.csv / val.csv / test.csv # tab-separated: <relative image path>\t<Malayalam word>
train/ val/ test/ # image files referenced by the corresponding CSV
Each CSV row maps one image file (path relative to its split folder) to its
ground-truth Malayalam word transcription.… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-ocr-words.malayalam-asr-corpus
Malayalam ASR Corpus
A multi-corpus Malayalam speech dataset aggregated from 5 public sources,
created for fine-tuning ASR models on Malayalam language.
This dataset was used to train
sajilck/whisper-small-malayalam
— the first multi-corpus Malayalam Whisper model on HuggingFace,
achieving 37.64% WER on CommonVoice 25 Malayalam test set.
Source Corpora
Corpus
Domain
Speaker Type
License
IMaSC
TTS / Read speech
Studio speakers
CC BY 4.0
SMC Malayalam… See the full description on the dataset page: https://huggingface.co/datasets/sajilck/malayalam-asr-corpus.arcastt-malayalam-dataset
arca-tuner -- Stream-A community (ml_cs)
Auto-generated by notebooks/gather_and_combine_datasets.ipynb.
profile: ml_cs
loaders: ['kathbath', 'indicvoices', 'shrutilipi', 'fleurs', 'imasc', 'indictts_ml', 'smc_msc', 'spring_inx', 'vaani', 'indicvoices_r']
samples: 533158
total hours: 1119.09
viewer-friendly parquet shards: data/train-*.parquet (246 shard(s); audio embedded as bytes)
training manifest: manifests/combined_streamA.jsonl (full schema, nested fields)
Licences are… See the full description on the dataset page: https://huggingface.co/datasets/taphuynh/arcastt-malayalam-dataset.pair_tamil_malayalam_ipa_transcription_romanizedmalayalam_common_voice_benchmarkingOCR-Bench1000-Malayalam
OCR-Bench1000-Malayalam
1000 synthetic printed-text line images with ground-truth transcriptions,
sampled from a larger locally-held Malayalam OCR training corpus.
This is a benchmark/sample release, not the full training set.
Data fields
Field
Description
file_name
relative path to the image (images/...)
text
ground-truth transcription
category
malayalam_only / english_only / mixed / numeric_and_symbols
length_bucket
short / medium / long, by… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Malayalam.malayalam_msc_benchmarkingpair_tamil_malayalam_ipa
Dataset Card for "pair_tamil_malayalam_ipa"
More Information needed
combined_malayalamindic-Malayalamhospitalcall_malayalammalayalam-orpheus-24khzmalayalam_wikiCommon Crawl - Malayalam.details_abhinand__malayalam-llama-7b-instruct-v0.1
Dataset Card for Evaluation run of abhinand/malayalam-llama-7b-instruct-v0.1
Dataset automatically created during the evaluation run of model abhinand/malayalam-llama-7b-instruct-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abhinand__malayalam-llama-7b-instruct-v0.1.Pronunciation-dictionary-malayalam
Malayalam Pronunciation Dictionary
This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here
It gives Phonemic transcription of Malayalam words in IPA format.
Dataset Details
Dataset Description
This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon]
(https://pypi.org/project/mlphon/) Python library.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.IndicTTS_Malayalam
Malayalam Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Malayalam monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Malayalam
Total Duration: ~17.89 hours (Male: 9.7 hours, Female: 8.19 hours)
Audio Format: WAV
Sampling… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Malayalam.HINDI-BENGALI-MALAYALAM-ODIA-VISUAL-GENOMESAM-LLAVA-20k-Malayalam-Caption-Pretrain
Malayalam translated version of unography/SAM-LLAVA-20k
Translated using indictrans2
Translation Code : code
malayalam-language-resources
Malayalam Language Resources
This CC-BY-4.0 release contains 21 JSONL catalogue and schema records for Malayalam, Sanskrit, English, translation, speech, and OCR collections. It is a template release: no third-party text, recordings, or scanned works are included.
Each future record must include its source, rights, consent status, and the licence that applies to that source.
indicvoices-v1amalayalam-speech-datasetMalayalam-VQA
Malayalam translated version of merve/vqav2-small
Translated using indictrans2
Translation Code : code
malayalam_speech
Malayalam Speech Dataset (VibeVoice)
This dataset contains Malayalam speech audio files and their corresponding transcriptions, prepared for fine-tuning ASR (Automatic Speech Recognition) models like VibeVoice.
Dataset Description
The dataset consists of approximately 4,126 Malayalam audio clips, split into training and testing sets with a 90/10 ratio, stratified by speaker gender.
Total Audio Files: 4,126
Train Split: 3,712 samples
Test Split: 414 samples
Total… See the full description on the dataset page: https://huggingface.co/datasets/ArjunJ/malayalam_speech.
