CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cognitive-Lab /Aya_Marathitabular1M<n<10M2 likes263 downloads3y agoHugging Face02kalpesh77 /marathi-phonology-matrices मराठी व्याकरण आणि ध्वनी मॅट्रिक्स Marathi Phonology Matrices गणितीय ध्वनी संश्लेषणासाठी (Mathematical Speech Synthesis) तयार केलेला सर्वसमावेशक मराठी फोनोलॉजी डेटासेट. 🎯 उद्देश्य हा डेटासेट मराठी भाषेच्या: फोनोलॉजिकल विश्लेषण मॉर्फोलॉजी (लिंग, वचन, काळ) संधि व श्व नियम युक्तक्षर (Clusters) Duration & Pitch नियम Loanword adaptation या सर्वांसाठी संरचित डेटा पुरवतो. TTS, ASR, G2P आणि Computational Linguistics संशोधनासाठी उपयुक्त. 📊… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/marathi-phonology-matrices.tabulartext-to-speech1K<n<10K1 likes100 downloads2mo agoHugging Face03shubhamugare /MMLU-Philosophy-Marathi MMLU Philosophy Questions in Marathi This dataset contains philosophy questions from the MMLU (Massive Multitask Language Understanding) benchmark translated into Marathi. Dataset Information Source: MMLU Philosophy subset from cais/mmlu Translation API: OpenAI GPT-4 Languages: English (original) and Marathi (translated) Total Questions: 311 Task Type: Multiple choice questions with 4 options each Dataset Structure Each row contains: original_question: The… See the full description on the dataset page: https://huggingface.co/datasets/shubhamugare/MMLU-Philosophy-Marathi.tabularquestion-answeringn<1K0 likes21 downloads1y agoHugging Face04InfoBayAI /Marathi-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Marathi STEM textbook data, containing 173 books and 7.81 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Marathi. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-STEM-Textbook-Dataset.tabular10K<n<100K0 likes19 downloads7d agoHugging Face05InfoBayAI /Marathi-Non-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Marathi Non-STEM textbook data, containing 584 books and 37.36 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Marathi. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ Billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-Non-STEM-Textbook-Dataset.tabular10K<n<100K1 likes15 downloads7d agoHugging Face06Reubencf /marathi-czech-sentences This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. marathi_czech_sentences This dataset contains short sentences and questions primarily in Marathi and Czech, covering various conversational contexts. The samples include inquiries about objects, actions, and origins, as well as exclamations and statements. It appears to be a multilingual collection focused on everyday dialogue structures. Dataset size There are 3… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/marathi-czech-sentences.tabular1K<n<10K0 likes13 downloads5mo agoHugging Face07SayantanJoker /audio_tts_description_marathitabular10K<n<100K0 likes11 downloads2y agoHugging Face08atx-labs /marathi-codemix-qagated Marathi Minglish QA ~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles. Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms. Example Question: Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi? Answer: Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.tabulartext-generation1M<n<10M0 likes9 downloads1mo agoHugging Face09MatrixSpeechAI /All_Marathi_ASR_stage_1tabular10K<n<100K0 likes7 downloads2y agoHugging Face10MatrixSpeechAI /All_Marathi_ASR_stage_2tabular10K<n<100K0 likes7 downloads2y agoHugging Face11rayman114 /All_Marathi_ASR_stage_3tabular10K<n<100K0 likes4 downloads7mo agoHugging Face12MatrixSpeechAI /All_Marathi_ASR_stage_3tabular10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.