datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-synthetic-profiles
🇮🇳 Indian Synthetic Identity Dataset
10,000 realistic Indian synthetic identities across 8 languages — generated by indic-faker
Dataset Description
This dataset contains 10,000 rows of realistic, synthetic Indian identity data generated using the indic-faker Python library. Every record is algorithmically valid — Aadhaar numbers pass Verhoeff checksum verification, GSTINs have correct state codes, and names are culturally authentic across 8 Indian languages.… See the full description on the dataset page: https://huggingface.co/datasets/adwaith06/indic-synthetic-profiles.profile-qa-synthetic-public-v1
Profile-QA Synthetic Public V1
Description
This dataset contains deterministic synthetic Q&A examples for public
resume/profile answering. It was generated from generic resume sections and
public-style facts, with evidence references back to section_id and fact_id.
The ontology is intentionally reusable across people and forks:
identity, current_role, experience, projects, education,
recommendations, skills, and interests. Temporal and practical sections
are… See the full description on the dataset page: https://huggingface.co/datasets/justinthelaw/profile-qa-synthetic-public-v1.persona-profiles-1m
