datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pronunciation-dictionary-malayalam
Malayalam Pronunciation Dictionary
This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here
It gives Phonemic transcription of Malayalam words in IPA format.
Dataset Details
Dataset Description
This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon]
(https://pypi.org/project/mlphon/) Python library.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.malayalam_news_classificationenglish-malayalam-names
English Malayalam names
This dataset has 27814162 person names both in English and Malayalam.
The source for this dataset is various election roles published by Government.
Potential usages:
English <-> Malayalam name transliteration tasks
Named entity recognition
Person name recognition
License
Creative commons Attribution Share Alike 4.0
Contact
Santhosh Thottingal santhosh.thottingal @ gmail.com
Malayalam-word-freq
Word Frequency Profile of Malayalam
The repo contains Malayalam words and their frequencies as obtained from AI4Bharat Indic NLP corpus.
There is an associated python script to plot the word frequnecy profile.
English_Malayalam_Translation_Human_annotated
English-Malayalam Government Parallel Corpus Synth
This dataset contains synthetic machine-translated English-Malayalam sentence pairs aligned from government and administrative text.
Machine Translation Notice
All parallel text in this dataset should be treated as synthetic machine-translated data. It is intended for research, corpus filtering, model adaptation, and experimentation. It should not be treated as human-verified gold translation without additional… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/English_Malayalam_Translation_Human_annotated.Ultimate-Malayalam-Dataset
About
This is a large dataset that contains a lot of malayalam text. The dataset is created by combining many other malayalam datasets, filtering and cleaning them.
The dataset is ideal to pretrain (or maybe even fine tune) a Large Language Model on Malayalam language.
This is the second version of the dataset where I've added some more data and removed a lot of data, which were very short(less than 75 characters). Having a lot off shorter sentences will significantly lower data… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/Ultimate-Malayalam-Dataset.Alpaca_Instruct_Malayalammalayalam-morphology-analyser
Malayalam Morphology Analyser dataset
This is a dataset of 801585 Malayalam words and their morphology analysis using Mlmorph Malayalam morphology analyser.
License
Creative Commons Attribution Share Alike 4.0
Contact
Contact santhosh.thottingal @ gmail.com
malayalam-sangrahaThis is a cleaned version of the malayalam subset of sangraha dataset. This only contains the human verified part of the dataset(which is high quality data obtained from Indic language PDFs, transcribed data from various Indic language videos, podcasts, movies, courses, etc.)
The csv dataset has around 6.3M rows, accounting to 32.8 GB. I've also removed the doc_id provided in the dataset, making this ideal for pretraining malayalam LLM.
For pretraining, I recommend using this dataset along… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/malayalam-sangraha.malayalam-visual-genome-instruction-setMalayalam-ArticlesMalayalam-Speech-Dataset
🎧 Malayalam Speech Dataset
The Malayalam Speech Dataset is a high-quality speech audio dataset designed to power AI and machine learning systems with reliable and diverse audio data. It includes 95 hours of recorded speech data across 650 files, available in MP3 and WAV formats, with a total size of 227 MB. This well-structured audio dataset delivers balanced and representative voice data, featuring 54% female and 46% male speakers, with age groups ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Malayalam-Speech-Dataset.Norm_Malayalam_Evaluation_samples
Malayalam ASR Reference Prediction dataset
This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset.
ASR Model Name: vrclc/Whisper_small_malayalam
Dataset: google/fleurs
Curated by: VRCLC
vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data.
The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model
The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Norm_Malayalam_Evaluation_samples.Flores-subsetbad_malayalam_datasetMalayalam_ner_taggedMalayalam_first_ready_for_sentimentMal_ASR_Predict_Ref_Samples
Malayalam ASR Reference Prediction dataset
This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset.
ASR Model Name: vrclc/Whisper_small_malayalam
Dataset: google/fleurs
Curated by: VRCLC
vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data.
The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model
The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Mal_ASR_Predict_Ref_Samples.malayalam-kannada-tamil-telugu-samam-dataset
Samam.net Multilingual Dictionary Dataset
Dataset Description
This dataset contains multilingual dictionary entries scraped from samam.net, a comprehensive South Indian language dictionary. The dataset provides translations between Malayalam and three other Dravidian languages: Kannada, Tamil, and Telugu.
Important Note about Script Usage
All text in this dataset is written in Malayalam script, even for non-Malayalam languages. This is a key characteristic of… See the full description on the dataset page: https://huggingface.co/datasets/cazzz307/malayalam-kannada-tamil-telugu-samam-dataset.colloquial-malayalamMalayalamMalayalamQAfinanceSheet1Malayalam_Dialects
