datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marathi-generated_4o-mini_2Mmarathi-land-law-reasoning
Marathi Land Legal Reasoning (100 High-Quality Rows)
📌 Dataset Description
This is a high-precision reasoning dataset focused on the Maharashtra Land Revenue Code (MLRC) and property succession laws in India. It is specifically designed for fine-tuning Large Language Models (LLMs) to handle complex, domain-specific logic in the Marathi language.
Unlike generic datasets, this collection focuses on "Chain-of-Thought" (CoT), providing a detailed thought block for every… See the full description on the dataset page: https://huggingface.co/datasets/aditya-datasets/marathi-land-law-reasoning.marathi-alpaca-llama-finetune
Marathi Alpaca Dataset for llama-finetune
This dataset contains 48,897 high-quality Marathi instruction-following examples, converted to the llama-finetune format.
Format
Each line in the JSONL file contains:
{
"messages": [
{
"role": "user",
"content": "निरोगी राहण्यासाठी तीन टिपा द्या."
},
{
"role": "assistant",
"content": "1. संतुलित आणि पौष्टिक आहार घ्या..."
}
]
}
Usage
Download
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/aghatage/marathi-alpaca-llama-finetune.adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-digital payments and Banking terms and topics- Hindi, Marathi, Bhojpuri, Maithili
This dataset contains question-and-answer pairs focused on personal finance and banking services in India, covering topics like UPI, net banking, tax payments, and government loan schemes. Each sample includes a user query followed by a detailed, step-by-step completion that provides actionable advice… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili.GPTeacher-MarathiIndicNLP-Marathimarathi-tinystories-10kmarathi-czech-sentences
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
marathi_czech_sentences
This dataset contains short sentences and questions primarily in Marathi and Czech, covering various conversational contexts. The samples include inquiries about objects, actions, and origins, as well as exclamations and statements. It appears to be a multilingual collection focused on everyday dialogue structures.
Dataset size
There are 3… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/marathi-czech-sentences.sample-science-marathialpaca-marathi-filteredmarathi-maharashtra-multidomain-SFT-1k
Marathi Maharashtra Multidomain SFT - 1K Sample
Dataset Description
This is a carefully curated 1,000-sample subset of the comprehensive Marathi-Maharashtra multidomain supervised fine-tuning (SFT) dataset. This high-quality dataset contains question-answer pairs covering diverse aspects of Marathi language, culture, history, and Maharashtra-related topics.
Key Features
High-Quality Human Verification: All responses have been verified by Marathi language… See the full description on the dataset page: https://huggingface.co/datasets/grpathak22/marathi-maharashtra-multidomain-SFT-1k.marathi-generated_4o-mini_2Mcolloquial-marathi-datasetmarathi-model-dataset
