datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us-names-by-state
US Baby names
The SSA dataset with baby names:
https://www.ssa.gov/OACT/babynames/
Coniferest
We use this dataset in the active anomaly discovery Python package coniferest:
https://coniferest.snad.space/en/latest/notebooks/us-names.html
Update the data
Install Python packages: pip install requests aiohttp universal_pathlib pandas
Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip
./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.russian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.d-info-2005-names
German Name Frequencies by State & District (D-Info 2005)
Regional frequency of surnames and forenames in Germany, from the D-Info
2005 telephone-directory CD-ROM (klickTel, data status 02.06.2005), at two
administrative levels aligned with census-2022 geography:
State = Bundesland — the 16 federal states.
District = Landkreis / kreisfreie Stadt — the 400 districts, keyed by
their 5-digit Kreisschlüssel (AGS).
For every name each table gives its number of 2005 telephone… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/d-info-2005-names.us_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/us_ssa_gender_neutral_first_names.train_names_imbalanced
WA Voter Names — unbalanced train split
Training split for binary name classification, built from the Washington State voter
registration database (VRDB) extract dated 2026-09-01. Natural class prevalence.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.
Related repo
Contents
Kymera-Solutions/train_names_balanced
same positives, negatives downsampled 1:1
test split
not yet uploaded — required for evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_imbalanced.ML-Music-Classifier-dataset-and-model-name-Models
🎧 Spotify Music Preference Analysis
🧠 Project Overview
This project analyzes Spotify music data to predict song preferences using machine learning models. The analysis is based on a dataset of 195 songs (100 liked, 95 disliked) with various audio features extracted from Spotify's API.
📂 Dataset Description
📥 Data Collection Process
Liked Songs (100 tracks):
🎵 Primarily French Rap
🎸 Some American Rap, Rock, and Electronic music
✅… See the full description on the dataset page: https://huggingface.co/datasets/Jack1808/ML-Music-Classifier-dataset-and-model-name-Models.tn-water-panels
Tamil Nadu Water Panels
Cleaned, analysis-ready hydrological series for the Cauvery basin and Tamil Nadu's
major reservoirs, assembled from Indian government open data.
Why this exists. The underlying data is public but not usable as published. The
national water portal's CWC daily reservoir dataset covers only Odisha and Madhya
Pradesh, and enumerating its Tamil Nadu resources returned no reservoir file. The
archived reservoir bulletins are weekly PDFs across two incompatible… See the full description on the dataset page: https://huggingface.co/datasets/nameissakthi/tn-water-panels.train_names_balanced
WA Voter Names — balanced train split
1:1 downsampled training split for binary name classification, built from the
Washington State voter registration database (VRDB) extract dated 2026-09-01.
Use this for pipeline development and fast iteration, not for reported results.
Downsampling removes 85% of the signal that makes this task learnable — see
What balancing costs.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_balanced.french_first_names_insee_2024
French First Names from Death Records (1970-2024)
This dataset contains French first names extracted from death records provided by INSEE (French National Institute of Statistics and Economic Studies) covering the period from 1970 to September 2024.
Dataset Description
Data Source
The data is sourced from INSEE's death records database. It includes first names of deceased individuals in France, providing valuable insights into naming patterns across different… See the full description on the dataset page: https://huggingface.co/datasets/eltorio/french_first_names_insee_2024.baby_namestitanicsynthetic_names_balanced_100kcanada_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/canada_ssa_gender_neutral_first_names.username-scarcity-and-handle-markets
NameSniper open datasets
First-party data on username scarcity and handle markets, published by NameSniper, a name checker and handle monitor. Study pages, methods and the latest figures live at https://namesniper.pro/research.
Everything here is free to reuse under CC BY 4.0: quote it, chart it, republish it, commercially or not. The one condition is a credit to NameSniper with a link to https://namesniper.pro/research or to the study you used.
Dataset
What it is
Period… See the full description on the dataset page: https://huggingface.co/datasets/NameSniperPro/username-scarcity-and-handle-markets.russian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names"… See the full description on the dataset page: https://huggingface.co/datasets/ssuverin/russian-names.nls_place_namesSource:
National Land Survey of Finland
Fields:
"placeNameId",
"spelling",
"language",
"parallelName_fin",
"parallelName_swe",
"placeType_fin"
"placeType_swe",
"placeType_eng",
"municipality_fin",
"municipality_swe",
"region",
"x",
"y"
Coordinate system:
EPSG:3067 (ETRS-TM35FIN)
Date created:
2026-07-23
russian-given-names-nen
Russian Given Names (NEN) — 1,551 names with meanings and ZAGS popularity
1,551 Russian given names (809 male, 742 female) with origin, short meaning, an editorial etymology note, diminutive and international forms, and popularity ranks among newborns based on open data from Moscow civil registry offices (ZAGS).
Curated by the editorial team of NEN («Нет, это нормально»), a Russian parenting magazine. Every record links to a full name page at n-e-n.ru/imena — the living catalog… See the full description on the dataset page: https://huggingface.co/datasets/MentalTech/russian-given-names-nen.france_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/france_ssa_gender_neutral_first_names.dataset_namename_finder_v1dataset_repository_name
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/namok21/dataset_repository_name.dataset_repository_name
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/egoing/dataset_repository_name.dataset_repository_name
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/thiefcat/dataset_repository_name.dataset_repository_name
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/ej94/dataset_repository_name.NameInfluence
NameInfluence
tags: names, influence mapping, social network analysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'NameInfluence' dataset comprises records detailing the influence of names within various social networks. It includes metrics derived from social media platforms, literature citations, and public recognition surveys. Each record contains the name of an individual or a character, the sources of their influence… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/NameInfluence.ap1datafraud_dataset
