datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Marathi-WikipediaAll_Marathi_ASRIndicTTS_Marathi
Marathi Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Marathi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Marathi
Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Marathi.Aya_Marathimarathi-tts-kathbath
Marathi Kathbath Speech Corpus
Dataset Summary
Marathi Kathbath Speech Corpus is a reformatted subset of the Kathbath speech collection focused on Marathi dialogue and sentence readings. Optimized for voice cloning, prosody modeling, and TTS acoustic model training.
Dataset Structure
Language: Marathi (mr)
Content: Thousands of Marathi sentence audio clips with normalized Devanagari text annotations.
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Upadhyay/marathi-tts-kathbath.marathi-speech-datasettask1086_pib_translation_marathi_english
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1086_pib_translation_marathi_english
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1086_pib_translation_marathi_english.OCR-Bench1000-Marathi
OCR-Bench1000-Marathi
1000 synthetic printed-text line images with ground-truth transcriptions,
sampled from a larger locally-held Marathi OCR training corpus.
This is a benchmark/sample release, not the full training set.
Data fields
Field
Description
file_name
relative path to the image (images/...)
text
ground-truth transcription
category
marathi_only / english_only / mixed / numeric_and_symbols
length_bucket
short / medium / long, by character… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Marathi.SPRING_INX_Marathi_R2task990_pib_translation_urdu_marathi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task990_pib_translation_urdu_marathi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task990_pib_translation_urdu_marathi.whisper-marathi-enTaken from the google/fleurs dataset (https://huggingface.co/datasets/google/fleurs)
Maps Marathi audio samples -> English text
Usage:
from datasets import load_dataset
dataset = load_dataset("Akchunks/whisper_marathi_en")
train_dataset = dataset["train"]
test_dataset = dataset["test"]
original_data_marathi_ttsmarathi-phonology-matrices
मराठी व्याकरण आणि ध्वनी मॅट्रिक्स
Marathi Phonology Matrices
गणितीय ध्वनी संश्लेषणासाठी (Mathematical Speech Synthesis) तयार केलेला सर्वसमावेशक मराठी फोनोलॉजी डेटासेट.
🎯 उद्देश्य
हा डेटासेट मराठी भाषेच्या:
फोनोलॉजिकल विश्लेषण
मॉर्फोलॉजी (लिंग, वचन, काळ)
संधि व श्व नियम
युक्तक्षर (Clusters)
Duration & Pitch नियम
Loanword adaptation
या सर्वांसाठी संरचित डेटा पुरवतो. TTS, ASR, G2P आणि Computational Linguistics संशोधनासाठी उपयुक्त.
📊… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/marathi-phonology-matrices.marathi_asr_dataset
Dataset Card for "marathi_asr_dataset"
More Information needed
marathi-english-cidco-long-docs
Marathi-English CIDCO Long Documents & Paragraphs Dataset
Parallel Marathi-to-English translation dataset extracted from CIDCO (City and Industrial Development Corporation of Maharashtra) government documents, resolutions, tenders, and town planning records.
This version is specifically filtered to retain only long, fluent sentences and multi-sentence paragraphs, removing single-word table cells, short headers, numbers, and form labels.
Dataset Splits
Split… See the full description on the dataset page: https://huggingface.co/datasets/anuj1541/marathi-english-cidco-long-docs.cidco-marathi-english-qa-scoredmarathi_english_datasetmarathi-alpaca-cleaned-translated
Marathi Alpaca Cleaned Translated
A Marathi translation of the 51,760-row Alpaca-Cleaned instruction-tuning dataset — Unsloth's hosted fork of yahma/alpaca-cleaned, which fixes hallucinations, empty outputs, and formatting errors found in the original Stanford Alpaca-52k dataset.
Translated using Meta's facebook/nllb-200-distilled-600M model. Built to reproduce and evaluate the Marathi instruction-tuning experiment from Khade et al., CHiPSAL 2025. The original paper translated… See the full description on the dataset page: https://huggingface.co/datasets/lubzo/marathi-alpaca-cleaned-translated.SPRING_INX_Marathi_R1marathi-english-cidco-docsmarathi-instruction-tuning-alpacaroopa-ca-2000-8-12-marathimarathi-english-cidco-long-docs-cleanedmarathi-pos-tagger
L3Cube-MahaPOS: Marathi Part-of-Speech Tagging Dataset
Dataset Description
L3Cube-MahaPOS is one of the first large-scale, manually annotated Part-of-Speech (POS) tagging datasets for Marathi — an Indo-Aryan language spoken by over 83 million people. The dataset comprises 32,354 sentences sourced from Marathi news text and annotated with a 16-tag scheme aligned with the Universal Dependencies (UD) v2 framework.
This dataset is part of the L3Cube-MahaNLP family of… See the full description on the dataset page: https://huggingface.co/datasets/l3cube-pune/marathi-pos-tagger.marathi-dictionary
Marathi Dictionary
Marathi-to-Marathi Dictionary
Dataset Description
Synthetically generated meanings in marathi for over 38k words
marathi-tts-indictts
Marathi IndicTTS Speech Corpus
Dataset Summary
Marathi IndicTTS Speech Corpus is a curated Marathi subset derived and formatted from IndicTTS studio recordings. It features clear, studio-recorded Marathi speech paired with phonetically balanced Devanagari transcriptions designed for acoustic feature extraction and neural vocoder training.
Dataset Structure
Format: High-fidelity WAV files (16-bit PCM, 48kHz / 22.05kHz) + normalized transcripts.… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Upadhyay/marathi-tts-indictts.cidco-marathi-english-blocks-combinedmarathi-tts-fleurs
Marathi FLEURS Speech Synthesis Dataset
Dataset Summary
Marathi FLEURS Speech Synthesis Dataset is a processed subset of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) benchmark for Marathi. The dataset is specially processed and formatted for Marathi TTS and ASR evaluation.
Dataset Structure
Language: Marathi (mr)
Script: Devanagari
Audio Spec: 16kHz MONO WAV audio.
Data Fields
id: Utterance… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Upadhyay/marathi-tts-fleurs.ICON26-COILD-INDIC-MT-Gujarati-Marathi
COILD-INDIC-MT 2026 — Gujarati–Marathi Dataset
This dataset is provided for the COILD-INDIC-MT 2026 Shared Task,
co-located with ICON 2026.
The shared task aims to foster research and innovation in
Natural Language Processing (NLP) for Indian Languages.
This repository contains data specifically for the:
Gujarati ↔ Marathi
language pair.
🔐 Access to the Dataset
This is a restricted and gated dataset.
Access is available only to authorized participants of the… See the full description on the dataset page: https://huggingface.co/datasets/ainlpml-iitp/ICON26-COILD-INDIC-MT-Gujarati-Marathi.english-to-marathi-mt
