datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indian-Constitution
Indian Constitution Dataset
The dataset can be used for text classification, text generation and text2text generation
oncology-financial-reasoning-india
🩺 Medzz-AI: Oncology & Financial Reasoning (India)
Status: Active | Context: Indian Healthcare | Focus: Clinical + Economic Logic
👋 The Problem: Why Current Medical AI Fails
State-of-the-art LLMs excel at clinical diagnosis but often fail at Health Economics. When asked to generate treatment plans, they frequently hallucinate costs, ignore local insurance constraints, or suggest financially viable treatments that are practically impossible for the patient.
Medzz-AI… See the full description on the dataset page: https://huggingface.co/datasets/Medzza/oncology-financial-reasoning-india.Travel_indiaindian-exam-corpus
Indian Exam Corpus
Overview
Indian Exam Corpus is an English educational text corpus designed for language model pretraining and educational NLP research.
The corpus consists of long-form educational documents covering topics commonly found in Indian competitive examinations.
Current subsets include:
JEE (Joint Entrance Examination)
NEET (National Eligibility cum Entrance Test)
Each document is stored as a single training example together with its associated… See the full description on the dataset page: https://huggingface.co/datasets/roshan-soni/indian-exam-corpus.
