datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MQAtky_persona_mqa
tky_persona_mqa Dataset Overview
tky_persona_mqa.json stores 67,575 narrated day-in-the-life summaries from Tokyo trajectories, each paired with two persona labels chosen by GPT-5.
Field Definitions
user_id: String linking the narrative back to its trajectory instance.
text: English-language narrative describing hourly activities inferred from GPS traces and nearby POIs.
choice: Two ordered persona labels. GPT-5 places its most plausible persona first and the least… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/tky_persona_mqa.vmlu-vi-mqa-answers-initial-phasenyc_persona_mqa
nyc_persona_mqa Dataset Overview
nyc_persona_mqa.json captures 65,115 narrated day-in-the-life summaries from New York City visitor and resident trajectories, each annotated with two persona hypotheses drafted by GPT-5 based on observed movement patterns and nearby POIs.
Field Definitions
user_id: Identifier that links the narrative back to a specific NYC trajectory instance.
text: English summary describing hourly activities inferred from GPS traces and contextual POI… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/nyc_persona_mqa.MQA
