datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aya_Marathimarathi-phonology-matrices
मराठी व्याकरण आणि ध्वनी मॅट्रिक्स
Marathi Phonology Matrices
गणितीय ध्वनी संश्लेषणासाठी (Mathematical Speech Synthesis) तयार केलेला सर्वसमावेशक मराठी फोनोलॉजी डेटासेट.
🎯 उद्देश्य
हा डेटासेट मराठी भाषेच्या:
फोनोलॉजिकल विश्लेषण
मॉर्फोलॉजी (लिंग, वचन, काळ)
संधि व श्व नियम
युक्तक्षर (Clusters)
Duration & Pitch नियम
Loanword adaptation
या सर्वांसाठी संरचित डेटा पुरवतो. TTS, ASR, G2P आणि Computational Linguistics संशोधनासाठी उपयुक्त.
📊… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/marathi-phonology-matrices.MMLU-Philosophy-Marathi
MMLU Philosophy Questions in Marathi
This dataset contains philosophy questions from the MMLU (Massive Multitask Language Understanding) benchmark translated into Marathi.
Dataset Information
Source: MMLU Philosophy subset from cais/mmlu
Translation API: OpenAI GPT-4
Languages: English (original) and Marathi (translated)
Total Questions: 311
Task Type: Multiple choice questions with 4 options each
Dataset Structure
Each row contains:
original_question: The… See the full description on the dataset page: https://huggingface.co/datasets/shubhamugare/MMLU-Philosophy-Marathi.Marathi-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Marathi STEM textbook data, containing 173 books and 7.81 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Marathi.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-STEM-Textbook-Dataset.Marathi-Non-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Marathi Non-STEM textbook data, containing 584 books and 37.36 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Marathi.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ Billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-Non-STEM-Textbook-Dataset.marathi-czech-sentences
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
marathi_czech_sentences
This dataset contains short sentences and questions primarily in Marathi and Czech, covering various conversational contexts. The samples include inquiries about objects, actions, and origins, as well as exclamations and statements. It appears to be a multilingual collection focused on everyday dialogue structures.
Dataset size
There are 3… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/marathi-czech-sentences.audio_tts_description_marathimarathi-codemix-qa
Marathi Minglish QA
~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles.
Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms.
Example
Question:
Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi?
Answer:
Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.All_Marathi_ASR_stage_1All_Marathi_ASR_stage_2All_Marathi_ASR_stage_3All_Marathi_ASR_stage_3
