datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arcastt-malayalam-dataset
arca-tuner -- Stream-A community (ml_cs)
Auto-generated by notebooks/gather_and_combine_datasets.ipynb.
profile: ml_cs
loaders: ['kathbath', 'indicvoices', 'shrutilipi', 'fleurs', 'imasc', 'indictts_ml', 'smc_msc', 'spring_inx', 'vaani', 'indicvoices_r']
samples: 533158
total hours: 1119.09
viewer-friendly parquet shards: data/train-*.parquet (246 shard(s); audio embedded as bytes)
training manifest: manifests/combined_streamA.jsonl (full schema, nested fields)
Licences are… See the full description on the dataset page: https://huggingface.co/datasets/taphuynh/arcastt-malayalam-dataset.malayalam_common_voice_benchmarkingmalayalam_msc_benchmarkingindicvoices-v1aspring_ml_conversationw2c-malayalamMalayalam-Non-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Malayalam Non-STEM textbook data, containing 149 books and 5.60 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Malayalam.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam-Non-STEM-Textbook-Dataset.audio_tts_description_malayalammalayalam-ocr-dataAya_Malayalam
Aya_Malayalam
This Dataset is curated from the original Aya-Collection dataset that was open-sourced by Cohere under the Apache-2.0 license.
The Aya Collection is a massive multilingual collection comprising 513 million instances of prompts and completions that cover a wide range of tasks. This collection uses instruction-style templates from fluent speakers and applies them to a curated list of datasets. It also includes translations of instruction-style datasets into 101… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/Aya_Malayalam.malayalam-kannada-tamil-telugu-samam-dataset
Samam.net Multilingual Dictionary Dataset
Dataset Description
This dataset contains multilingual dictionary entries scraped from samam.net, a comprehensive South Indian language dictionary. The dataset provides translations between Malayalam and three other Dravidian languages: Kannada, Tamil, and Telugu.
Important Note about Script Usage
All text in this dataset is written in Malayalam script, even for non-Malayalam languages. This is a key characteristic of… See the full description on the dataset page: https://huggingface.co/datasets/cazzz307/malayalam-kannada-tamil-telugu-samam-dataset.
