lbakar/health-log-extraction-dataset
To create this dataset, we sampled demographic seeds from an occupation and age-range table from the Labor Force Statistics[1]. For each occupation, the pipeline randomly selected an age range, sampled an age within that range, and assigned a gender from a fixed set of options. These demographic seeds were used to prompt an LLM to generate structured personas containing a name, description, medications or supplements, general mood, and possible health conditions or injuries. We then used… See the full description on the dataset page: https://huggingface.co/datasets/lbakar/health-log-extraction-dataset.
028
