datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicSafe
IndicSafe
Authors: Priyaranjan Pattnayak, Garima Panwar, and Sanchari Chowdhuri.
IndicSafe is a multilingual benchmark for evaluating large-language-model safety behavior across 12 South Asian languages. It contains 6,000 translated prompt rows: 500 source rows in each language, spanning harmful, harmless-control, and deliberately ambiguous categories.
Content warning: the benchmark contains prompts about hate, discrimination, violence, misinformation, political manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ppattnay/IndicSafe.Indic-post-ocr-correction
Indic Contextual Post-OCR Correction
Dataset Summary
This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of:
an OCR-generated sentence (noisy),
the preceding sentence used as context, and
the corrected sentence (ground truth).
Hugging Face dataset page:
https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction
Supported Tasks
Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.truthfulqa_indicOriginal Repository
Tasks (from original repository)
Generation (main task):
Task: Given a question, generate a 1-2 sentence answer.
Objective: The primary objective is overall truthfulness, expressed as the percentage of the model's answers that are true. Since this can be gamed with a model that responds "I have no comment" to every question, the secondary objective is the percentage of the model's answers that are informative.
Future Work:
Validate… See the full description on the dataset page: https://huggingface.co/datasets/vakyansh/truthfulqa_indic.indic-synthetic-profiles
🇮🇳 Indian Synthetic Identity Dataset
10,000 realistic Indian synthetic identities across 8 languages — generated by indic-faker
Dataset Description
This dataset contains 10,000 rows of realistic, synthetic Indian identity data generated using the indic-faker Python library. Every record is algorithmically valid — Aadhaar numbers pass Verhoeff checksum verification, GSTINs have correct state codes, and names are culturally authentic across 8 Indian languages.… See the full description on the dataset page: https://huggingface.co/datasets/adwaith06/indic-synthetic-profiles.
