CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VishnuPJ /Malayalam_CultureX_IndicCorp_SMCMalayalam Pretraining/Tokenization dataset. Preprocessed and combined data from the following links, * ai4bharat * CulturaX * Swathanthra Malayalam Computing Commands used for preprocessing. To remove all non Malayalam characters. sed -i 's/[^ം-ൿ.,;:@$%+&?!() ]//g' test.txt To merge all the text files in a particular Directory(Sub-Directory) find SMC -type f -name '*.txt' -exec cat {} ; >> combined_SMC.txt To remove all lines with characters less than 5. grep -P… See the full description on the dataset page: https://huggingface.co/datasets/VishnuPJ/Malayalam_CultureX_IndicCorp_SMC.text10M<n<100M4 likes1.4k downloads2y agoHugging Face02aoxo /asr_malayalamaudio1 likes681 downloads2y agoHugging Face03ussooraj /OCR-bench-Malayalamimageimage-to-text1K<n<10K0 likes637 downloads1mo agoHugging Face04rajeshradhakrishnan /malayalam_2020_wiki��This dataset is from the common-crawl-malayalam repo: https://github.com/qburst/common-crawl-malayalam 1 likes617 downloads4y agoHugging Face05Praha-Labs /indic-Malayalam-PDaudio10K<n<100K0 likes609 downloads1y agoHugging Face06SPRINGLab /SPRING_INX_Malayalam_R1audio10K<n<100K0 likes564 downloads2y agoHugging Face07psk /malayalam-speech-178h Malayalam Speech — 179 hours (denoised, unlabeled) 86,799 Malayalam speech clips, 48 kHz stereo WAV, ~7.4 s average. No transcripts — this is unlabeled audio, intended for self-supervised pretraining, voice/speaker modelling, or as raw material for your own labelling pipeline. Processing Each clip passed through a full source-separation and enhancement chain: Vocal isolation — BS-RoFormer De-reverberation — UVR-DeEcho-DeReverb Noise removal — UVR-DeNoise-Lite… See the full description on the dataset page: https://huggingface.co/datasets/psk/malayalam-speech-178h.audio10K<n<100K4 likes422 downloads2mo agoHugging Face08icfoss /malayalam-ocr-words Malayalam OCR Words A word-level Malayalam OCR dataset: cropped word images paired with their transcribed text label, split into train/validation/test sets. Dataset structure train.csv / val.csv / test.csv # tab-separated: <relative image path>\t<Malayalam word> train/ val/ test/ # image files referenced by the corresponding CSV Each CSV row maps one image file (path relative to its split folder) to its ground-truth Malayalam word transcription.… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-ocr-words.imageimage-to-text10K<n<100K0 likes340 downloads7d agoHugging Face09sajilck /malayalam-asr-corpus Malayalam ASR Corpus A multi-corpus Malayalam speech dataset aggregated from 5 public sources, created for fine-tuning ASR models on Malayalam language. This dataset was used to train sajilck/whisper-small-malayalam — the first multi-corpus Malayalam Whisper model on HuggingFace, achieving 37.64% WER on CommonVoice 25 Malayalam test set. Source Corpora Corpus Domain Speaker Type License IMaSC TTS / Read speech Studio speakers CC BY 4.0 SMC Malayalam… See the full description on the dataset page: https://huggingface.co/datasets/sajilck/malayalam-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes313 downloads2mo agoHugging Face10taphuynh /arcastt-malayalam-dataset arca-tuner -- Stream-A community (ml_cs) Auto-generated by notebooks/gather_and_combine_datasets.ipynb. profile: ml_cs loaders: ['kathbath', 'indicvoices', 'shrutilipi', 'fleurs', 'imasc', 'indictts_ml', 'smc_msc', 'spring_inx', 'vaani', 'indicvoices_r'] samples: 533158 total hours: 1119.09 viewer-friendly parquet shards: data/train-*.parquet (246 shard(s); audio embedded as bytes) training manifest: manifests/combined_streamA.jsonl (full schema, nested fields) Licences are… See the full description on the dataset page: https://huggingface.co/datasets/taphuynh/arcastt-malayalam-dataset.tabular100K<n<1M0 likes296 downloads5mo agoHugging Face11mugezhang /pair_tamil_malayalam_ipa_transcription_romanizedtext10M<n<100M0 likes277 downloads9mo agoHugging Face12kurianbenoy /malayalam_common_voice_benchmarkingtabular1K<n<10K1 likes239 downloads3y agoHugging Face13meharuhanzz /OCR-Bench1000-Malayalam OCR-Bench1000-Malayalam 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Malayalam OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category malayalam_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Malayalam.imageimage-to-text1K<n<10K0 likes233 downloads9d agoHugging Face14kurianbenoy /malayalam_msc_benchmarkingtabular10K<n<100K1 likes228 downloads3y agoHugging Face15krishnAbadikelA /pair_tamil_malayalam_ipa Dataset Card for "pair_tamil_malayalam_ipa" More Information needed text1M<n<10M0 likes222 downloads1y agoHugging Face16asr-malayalam /combined_malayalamaudio10K<n<100K0 likes213 downloads2y agoHugging Face17Praha-Labs /indic-Malayalamaudio10K<n<100K0 likes211 downloads1y agoHugging Face18anjalikrishna07 /hospitalcall_malayalamaudio1K<n<10K0 likes201 downloads2mo agoHugging Face19vysakh25 /malayalam-orpheus-24khzaudio10K<n<100K0 likes181 downloads7mo agoHugging Face20rajeshradhakrishnan /malayalam_wikiCommon Crawl - Malayalam.texttext-generation1M<n<10M2 likes171 downloads4y agoHugging Face21open-llm-leaderboard-old /details_abhinand__malayalam-llama-7b-instruct-v0.1 Dataset Card for Evaluation run of abhinand/malayalam-llama-7b-instruct-v0.1 Dataset automatically created during the evaluation run of model abhinand/malayalam-llama-7b-instruct-v0.1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abhinand__malayalam-llama-7b-instruct-v0.1.0 likes149 downloads3y agoHugging Face22kavyamanohar /Pronunciation-dictionary-malayalam Malayalam Pronunciation Dictionary This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here It gives Phonemic transcription of Malayalam words in IPA format. Dataset Details Dataset Description This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon] (https://pypi.org/project/mlphon/) Python library. Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.text100K<n<1M2 likes149 downloads2y agoHugging Face23SPRINGLab /IndicTTS_Malayalam Malayalam Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Malayalam monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Malayalam Total Duration: ~17.89 hours (Male: 9.7 hours, Female: 8.19 hours) Audio Format: WAV Sampling… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Malayalam.audiotext-to-speech10K<n<100K2 likes148 downloads2y agoHugging Face24ShivRamSaud /HINDI-BENGALI-MALAYALAM-ODIA-VISUAL-GENOMEimage10K<n<100K0 likes137 downloads18d agoHugging Face25VishnuPJ /SAM-LLAVA-20k-Malayalam-Caption-Pretrain Malayalam translated version of unography/SAM-LLAVA-20k Translated using indictrans2 Translation Code : code image10K<n<100K0 likes123 downloads2y agoHugging Face26kaumudi-ai /malayalam-language-resources Malayalam Language Resources This CC-BY-4.0 release contains 21 JSONL catalogue and schema records for Malayalam, Sanskrit, English, translation, speech, and OCR collections. It is a template release: no third-party text, recordings, or scanned works are included. Each future record must include its source, rights, consent status, and the licence that applies to that source. 0 likes116 downloads3mo agoHugging Face27asr-malayalam /indicvoices-v1atabular10K<n<100K0 likes115 downloads2y agoHugging Face28shunyalabs /malayalam-speech-datasetaudio100K<n<1M1 likes111 downloads1y agoHugging Face29VishnuPJ /Malayalam-VQA Malayalam translated version of merve/vqav2-small Translated using indictrans2 Translation Code : code image10K<n<100K0 likes101 downloads2y agoHugging Face30ArjunJ /malayalam_speech Malayalam Speech Dataset (VibeVoice) This dataset contains Malayalam speech audio files and their corresponding transcriptions, prepared for fine-tuning ASR (Automatic Speech Recognition) models like VibeVoice. Dataset Description The dataset consists of approximately 4,126 Malayalam audio clips, split into training and testing sets with a 90/10 ratio, stratified by speaker gender. Total Audio Files: 4,126 Train Split: 3,712 samples Test Split: 414 samples Total… See the full description on the dataset page: https://huggingface.co/datasets/ArjunJ/malayalam_speech.audioautomatic-speech-recognition1K<n<10K0 likes92 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.