CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kavyamanohar /Pronunciation-dictionary-malayalam Malayalam Pronunciation Dictionary This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here It gives Phonemic transcription of Malayalam words in IPA format. Dataset Details Dataset Description This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon] (https://pypi.org/project/mlphon/) Python library. Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.text100K<n<1M2 likes149 downloads2y agoHugging Face02mlexplorer008 /malayalam_news_classificationtext1K<n<10K0 likes88 downloads2y agoHugging Face03santhosh /english-malayalam-names English Malayalam names This dataset has 27814162 person names both in English and Malayalam. The source for this dataset is various election roles published by Government. Potential usages: English <-> Malayalam name transliteration tasks Named entity recognition Person name recognition License Creative commons Attribution Share Alike 4.0 Contact Santhosh Thottingal santhosh.thottingal @ gmail.com text10M<n<100M6 likes57 downloads3y agoHugging Face04kavyamanohar /Malayalam-word-freq Word Frequency Profile of Malayalam The repo contains Malayalam words and their frequencies as obtained from AI4Bharat Indic NLP corpus. There is an associated python script to plot the word frequnecy profile. image1M<n<10M0 likes41 downloads3y agoHugging Face05icfoss /English_Malayalam_Translation_Human_annotated English-Malayalam Government Parallel Corpus Synth This dataset contains synthetic machine-translated English-Malayalam sentence pairs aligned from government and administrative text. Machine Translation Notice All parallel text in this dataset should be treated as synthetic machine-translated data. It is intended for research, corpus filtering, model adaptation, and experimentation. It should not be treated as human-verified gold translation without additional… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/English_Malayalam_Translation_Human_annotated.texttranslation10K<n<100K0 likes34 downloads10d agoHugging Face06Arjun-G-Ravi /Ultimate-Malayalam-Dataset About This is a large dataset that contains a lot of malayalam text. The dataset is created by combining many other malayalam datasets, filtering and cleaning them. The dataset is ideal to pretrain (or maybe even fine tune) a Large Language Model on Malayalam language. This is the second version of the dataset where I've added some more data and removed a lot of data, which were very short(less than 75 characters). Having a lot off shorter sentences will significantly lower data… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/Ultimate-Malayalam-Dataset.texttext-generation1M<n<10M2 likes33 downloads10mo agoHugging Face07VishnuPJ /Alpaca_Instruct_Malayalamtext10K<n<100K6 likes31 downloads2y agoHugging Face08santhosh /malayalam-morphology-analyser Malayalam Morphology Analyser dataset This is a dataset of 801585 Malayalam words and their morphology analysis using Mlmorph Malayalam morphology analyser. License Creative Commons Attribution Share Alike 4.0 Contact Contact santhosh.thottingal @ gmail.com text100K<n<1M2 likes27 downloads3y agoHugging Face09Arjun-G-Ravi /malayalam-sangrahaThis is a cleaned version of the malayalam subset of sangraha dataset. This only contains the human verified part of the dataset(which is high quality data obtained from Indic language PDFs, transcribed data from various Indic language videos, podcasts, movies, courses, etc.) The csv dataset has around 6.3M rows, accounting to 32.8 GB. I've also removed the doc_id provided in the dataset, making this ideal for pretraining malayalam LLM. For pretraining, I recommend using this dataset along… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/malayalam-sangraha.texttext-generation1M<n<10M1 likes26 downloads10mo agoHugging Face10DebasishDhal99 /malayalam-visual-genome-instruction-settext10K<n<100K0 likes23 downloads11mo agoHugging Face11lkarjun /Malayalam-Articlestext10K<n<100K1 likes20 downloads5y agoHugging Face12Speech-data /Malayalam-Speech-Dataset 🎧 Malayalam Speech Dataset The Malayalam Speech Dataset is a high-quality speech audio dataset designed to power AI and machine learning systems with reliable and diverse audio data. It includes 95 hours of recorded speech data across 650 files, available in MP3 and WAV formats, with a total size of 227 MB. This well-structured audio dataset delivers balanced and representative voice data, featuring 54% female and 46% male speakers, with age groups ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Malayalam-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes14 downloads6mo agoHugging Face13asr-malayalam /Norm_Malayalam_Evaluation_samples Malayalam ASR Reference Prediction dataset This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset. ASR Model Name: vrclc/Whisper_small_malayalam Dataset: google/fleurs Curated by: VRCLC vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data. The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Norm_Malayalam_Evaluation_samples.textsentence-similarityn<1K0 likes10 downloads2y agoHugging Face14asr-malayalam /Flores-subsettext1K<n<10K0 likes9 downloads2y agoHugging Face15animaRegem /bad_malayalam_datasettext1M<n<10M0 likes6 downloads2y agoHugging Face16Achuth7Achu /Malayalam_ner_taggedtext10K<n<100K1 likes4 downloads2y agoHugging Face17wlkla /Malayalam_first_ready_for_sentimenttext1K<n<10K0 likes4 downloads2y agoHugging Face18asr-malayalam /Mal_ASR_Predict_Ref_Samples Malayalam ASR Reference Prediction dataset This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset. ASR Model Name: vrclc/Whisper_small_malayalam Dataset: google/fleurs Curated by: VRCLC vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data. The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Mal_ASR_Predict_Ref_Samples.textsentence-similarityn<1K0 likes3 downloads2y agoHugging Face19cazzz307 /malayalam-kannada-tamil-telugu-samam-datasetgated Samam.net Multilingual Dictionary Dataset Dataset Description This dataset contains multilingual dictionary entries scraped from samam.net, a comprehensive South Indian language dictionary. The dataset provides translations between Malayalam and three other Dravidian languages: Kannada, Tamil, and Telugu. Important Note about Script Usage All text in this dataset is written in Malayalam script, even for non-Malayalam languages. This is a key characteristic of… See the full description on the dataset page: https://huggingface.co/datasets/cazzz307/malayalam-kannada-tamil-telugu-samam-dataset.tabular10K<n<100K0 likes3 downloads1y agoHugging Face20MahishreeMajji24 /colloquial-malayalamtextn<1K0 likes2 downloads2y agoHugging Face21SnehaJay7 /Malayalamtext1K<n<10K0 likes2 downloads2y agoHugging Face22Vineeth12 /MalayalamQAfinanceSheet1textn<1K0 likes1 downloads2y agoHugging Face23Aby003 /Malayalam_Dialectstextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.