CoolFace
Datasetpublic

dominicDK94/nemotron-personas-lite

Nemotron-Personas Lite (10 countries × 10k) A ~63MB derived subsample of NVIDIA's Nemotron-Personas synthetic persona datasets (10 countries, ~24GB total in the originals), built for the persona-lightsim harness — lightweight persona market research and simulation with coding agents. What was derived, exactly Per country: 10,000 personas sampled with fixed seed 42 (shard-size-proportional, row-group random extraction) from the original train splits. Columns: 15… See the full description on the dataset page: https://huggingface.co/datasets/dominicDK94/nemotron-personas-lite.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes274downloads
Dataset Card

Nemotron-Personas Lite (10 countries × 10k)

A ~63MB derived subsample of NVIDIA's Nemotron-Personas synthetic persona datasets (10 countries, ~24GB total in the originals), built for the persona-lightsim harness — lightweight persona market research and simulation with coding agents.

What was derived, exactly

Per country: 10,000 personas sampled with fixed seed 42 (shard-size-proportional, row-group random extraction) from the original train splits.

  • —Columns: 15 of the original 26 — persona, professional_persona, arts_persona, hobbies_and_interests_list, skills_and_expertise_list, career_goals_and_ambitions, sex, age, marital_status, education_level, occupation, province, district, country, cultural_background
  • —Trimming: persona/professional_persona/arts_persona truncated to 400 chars, career_goals_and_ambitions/cultural_background to 300 chars (marked with a trailing …)
  • —Structure preserved: original shard file layout, including Belgium's language-quota shards (nl/fr/de/en_BE.parquet) and India's English-only shards (en_IN-*), so samplers written for the originals run unchanged
  • —Build script: `scripts/build_lite_pack.py`; per-file sha256 in `scripts/data_manifest.json`

Countries: Belgium, Brazil, El Salvador, France, India, Japan, Korea, Singapore, USA, Vietnam.

If you need full narratives or all 26 columns, use the NVIDIA originals instead.

Attribution / License

This dataset is a derivative work of the Nemotron-Personas dataset family by NVIDIA Corporation, each released under CC-BY-4.0 (Nemotron-Personas-Belgium, -Brazil, -El-Salvador, -France, -India, -Japan, -Korea, -Singapore, -USA, -Vietnam, produced with NeMo Data Designer). Changes made: column selection (15/26), narrative-field truncation, and per-country subsampling to 10k rows as described above. This derivative is likewise distributed under CC-BY-4.0.

The personas are synthetic — they represent statistically grounded population composition, not real individuals or behavioral logs.