indian-names
indian-names-1.5M
IndicNames-1.5M — Indian Names Dataset
A large-scale curated dataset of 1.5 million unique Indian names collected from multiple Indic languages and regions.
Suitable for:
Name generation (LLMs / GPT fine-tuning)
NLP experimentation with Indian names data
Tokenization / language model benchmarking
Dataset Structure
Each line contains one cleaned name:
indian_names
indian_names
Unique customer names (lowercased) extracted from call/loan record databases
(xPertVoice Aug/Sep, both "PQ" and "PQ2" variants), each paired with the
language(s) detected from s3_key.
customer_names_with_language.csv — two columns: name, language.
One row per unique name. If a name was seen under multiple languages,
all of them are listed in one comma-separated field (e.g. "english, tamil").
Language is detected by scanning s3_key for a known Indian-language… See the full description on the dataset page: https://huggingface.co/datasets/MeghanaKap/indian_names.Indian_Names_with_Gender_Dataset
🇮🇳 Indian Names & Gender Dataset (Balanced)
Dataset Summary
This dataset contains 42,000 samples designed for training models to identify Indian names and classify their gender. It is perfectly balanced across three categories, making it ideal for training robust classifiers that can distinguish between real names and random text.
Dataset Structure
Column
Type
Description
Name
String
The text string (Name or Random Word).
Label
Integer
Class ID… See the full description on the dataset page: https://huggingface.co/datasets/shisha-07/Indian_Names_with_Gender_Dataset.indian-states-person-namesindian-names-asr
