CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taphuynh /arcastt-malayalam-dataset arca-tuner -- Stream-A community (ml_cs) Auto-generated by notebooks/gather_and_combine_datasets.ipynb. profile: ml_cs loaders: ['kathbath', 'indicvoices', 'shrutilipi', 'fleurs', 'imasc', 'indictts_ml', 'smc_msc', 'spring_inx', 'vaani', 'indicvoices_r'] samples: 533158 total hours: 1119.09 viewer-friendly parquet shards: data/train-*.parquet (246 shard(s); audio embedded as bytes) training manifest: manifests/combined_streamA.jsonl (full schema, nested fields) Licences are… See the full description on the dataset page: https://huggingface.co/datasets/taphuynh/arcastt-malayalam-dataset.tabular100K<n<1M0 likes296 downloads5mo agoHugging Face02kurianbenoy /malayalam_common_voice_benchmarkingtabular1K<n<10K1 likes239 downloads3y agoHugging Face03kurianbenoy /malayalam_msc_benchmarkingtabular10K<n<100K1 likes228 downloads3y agoHugging Face04asr-malayalam /indicvoices-v1atabular10K<n<100K0 likes115 downloads2y agoHugging Face05asr-malayalam /spring_ml_conversationtabular10K<n<100K0 likes28 downloads2y agoHugging Face06nebulatgs /w2c-malayalamtabular10K<n<100K0 likes20 downloads3y agoHugging Face07InfoBayAI /Malayalam-Non-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Malayalam Non-STEM textbook data, containing 149 books and 5.60 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Malayalam. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam-Non-STEM-Textbook-Dataset.tabular10K<n<100K0 likes12 downloads8d agoHugging Face08SayantanJoker /audio_tts_description_malayalamtabular10K<n<100K0 likes10 downloads2y agoHugging Face09OdiaGenAIOCR /malayalam-ocr-datatabularn<1K0 likes8 downloads3mo agoHugging Face10Cognitive-Lab /Aya_Malayalamgated Aya_Malayalam This Dataset is curated from the original Aya-Collection dataset that was open-sourced by Cohere under the Apache-2.0 license. The Aya Collection is a massive multilingual collection comprising 513 million instances of prompts and completions that cover a wide range of tasks. This collection uses instruction-style templates from fluent speakers and applies them to a curated list of datasets. It also includes translations of instruction-style datasets into 101… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/Aya_Malayalam.tabular1M<n<10M1 likes4 downloads3y agoHugging Face11cazzz307 /malayalam-kannada-tamil-telugu-samam-datasetgated Samam.net Multilingual Dictionary Dataset Dataset Description This dataset contains multilingual dictionary entries scraped from samam.net, a comprehensive South Indian language dictionary. The dataset provides translations between Malayalam and three other Dravidian languages: Kannada, Tamil, and Telugu. Important Note about Script Usage All text in this dataset is written in Malayalam script, even for non-Malayalam languages. This is a key characteristic of… See the full description on the dataset page: https://huggingface.co/datasets/cazzz307/malayalam-kannada-tamil-telugu-samam-dataset.tabular10K<n<100K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.