datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.Nemotron-Personas-Japan
Nemotron-Personas-Japan
現実世界の分布に基づいたペルソナ生成のための複合AIアプローチ
データセット概要 (Dataset Overview)
Nemotron-Personas-Japan は、日本における人口の多様性と豊かさを捉えることを目的とし、実世界の人口統計、地理的分布、性格特性の分布に基づいて合成的に生成されたペルソナのオープンソースデータセットです。名前、性別、年齢、背景、婚姻状況、学歴、職業、居住地などの統計に基づいて生成した初のデータセットされた Nemotron-Personas の日本語版です。本バージョンでは、日本語における多様なモデリングユースケースに適した高品質のペルソナを提供します
Nemotron-Personas-Japan は、日本のモデル開発者が重要な地域固有の人口統計や文化的背景を取り入れたソブリンAIシステムを開発することを支援します。本データセットは、日本の地理的・人口統計的な実分布を反映することで、合成データの多様性を高め、バイアスを軽減し、model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Japan.Nemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and richness of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-India.multicam_no_personcoco-2017-person-val
Dataset Card for coco-2017-validation
This is a FiftyOne dataset with 2693 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("naufalso/coco-2017-person-val")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/coco-2017-person-val.Nemotron-Personas-France
Nemotron-Personas-France
Une approche d'IA composée pour des personas ancrés dans des distributions réelles
A compound AI approach to personas grounded in real-world distributions
Vue d'ensemble du jeu de données (Dataset Overview)
Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.Persona-MME
PersonaVLM: Long-Term Personalized Multimodal LLMs (CVPR 2026)
🎉 News: Our paper "PersonaVLM: Long-Term Personalized Multimodal LLMs" is accepted to CVPR 2026!
🌟 Introduction
PersonaVLM is an innovative personalized multimodal agent framework designed for long-term personalization. It transforms a general-purpose MLLM into a personalized assistant by integrating three key capabilities:
Remembering: Proactively extracts and summarizes… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Persona-MME.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.persona-fas-bench-v1
Filtered copy. Redistribution of saatvikbilla1/persona-fas with a small number of
images withheld and capture metadata (EXIF/GPS, XMP, IPTC) removed.
Attack images in this release: 21,056.
person-face-package-home-security-detection
Person, face & package — home-security detection dataset
YOLO-format dataset for person, face, vehicles, small vehicles, parcels, pets, and birds in home / delivery / street scenes. Full data lives under bigsplit/; sample/ is a 100-image preview subset (same layout: flat images/ and labels/).
Training configs point at a data.yaml beside those folders. Both splits use the same class list (nc and names); only the root path and which image set is packaged differ. Labels are YOLO .txt… See the full description on the dataset page: https://huggingface.co/datasets/OliseNS/person-face-package-home-security-detection.Thermal-Person-Detector
Dataset Card for ThermalPersonDetector
A thermal image dataset for detecting people in a scene. The dataset contains only one class person
This is a FiftyOne dataset with 8778 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Thermal-Person-Detector.change_eye_face_head_person
Overview
This data set contains manually curated, high quality images that can be used to
train image editing AI models like
FLUX.1 Kontext
to be able to take an input image and a reference image to create a target
image that is looking like the input image but with one of those parts replaced:
eyes
face
head
person (input image cloths are kept)
person (reference image cloths are kept)
Typical prompts for this editing could then be:
Change the eyes, keeping the rest of the image… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/change_eye_face_head_person.Nemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/Nemotron-Personas-India.Persona-Fluxed-10k-2608
Persona Fluxed 10k
Synthetic persona portraits rendered with FLUX.2-klein-4b (8-step, 1024x1024) from the NVIDIA Nemotron-Personas-* datasets.
Each persona is grounded in real-world demographic, geographic and
personality-trait distributions for its country (CC BY 4.0 source; no real people).
Currently Nemotron-Personas exist for:
USA — English
Japan — Japanese
India — English, Hindi
Brazil — Portuguese
Singapore — English
France — French
Korea — Korean
El Salvador — Spanish… See the full description on the dataset page: https://huggingface.co/datasets/retowyss/Persona-Fluxed-10k-2608.MegaDepth-v1this-person-does-not-exist-10kpersona-fas-previewperson-centric-images-stable-diffusion-v1-1PersonaChat-Qwen-Image-2512-enhancedPerson_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform
Person Detection and Re-Identification from Low Altitude UAV-based Platform
Dataset Description
This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format.
The dataset supports two tasks:
Person Detection — detecting people in aerial drone footage
Person Re-Identification (Re-ID) — recognizing and… See the full description on the dataset page: https://huggingface.co/datasets/Mikiee/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.PersonaChat-Qwen-Image-2512-originalSynthetic-Persona-Chat-Qwen-Image-2512-originalperson-background-dataset
Person-Background Dataset
Diverse people placed in various scene backgrounds. Generated with FLUX.1-dev (people) and FLUX.1-Kontext (background editing).
How It Is Collected
The collect.py script:
Stage 1 – People: Generates 10 diverse people with FLUX.1-dev (different ethnicities, genders, ages).
Stage 2 – Backgrounds: Uses FLUX.1-Kontext img2img to edit each person into scene categories (forest, beach, office, etc.). Preserves person identity while changing only the… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/person-background-dataset.inria-personshrec_empathic
Social Human Robot Embodied Conversation (SHREC) Dataset: Empathic Subset (RSS 2026)
The SHREC Empathic subset contains real-world human-robot interaction video data from Shen et al. (2024), collected over a month-long deployment of social robots in participants’ homes, as participants engage in natural, empathic storytelling interactions with AI agents.
Authors: Dong Won Lee, Yubin Kim, Sooyeon Jeong, Denison Guvenoz, Parker Malachowsky, Louis-Philippe Morency, Cynthia… See the full description on the dataset page: https://huggingface.co/datasets/MIT-personal-robots/shrec_empathic.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/neuralnetworker/Nemotron-Personas-Korea.psychometric-fa-runsinria-personpersona-faswx_hand_drawn_personae_style
微信手绘人物风格素材
本数据集包含两组手绘人物/角色风格的原像素 PNG 分切素材,共 32 张图片。
数据结构
data/
└── train/
├── persona_set_01/ # 第一组人物素材,16 张
└── happy_dog/ # 快乐修狗素材,16 张
仓库根目录还保留了两组原始上传目录及对应的 裁切信息.txt,用于追溯原始分切参数。
数据字段
image:原像素 PNG 图片。
本数据集不生成分类标签,加载结果仅包含 image 字段。
使用方式
from datasets import load_dataset
dataset = load_dataset("hi-syzh/wx_hand_drawn_personae_style")
print(dataset["train"][0])
素材说明
目录
数量
说明
persona_set_01… See the full description on the dataset page: https://huggingface.co/datasets/hi-syzh/wx_hand_drawn_personae_style.
