datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atomic-metrics-demographic-training-size
Atomic Metrics: Demographic Training-Size Analysis
Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction.
Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.synthetic_demographics_seed
Synthetic Demographic Seeds v1
This is a dataset of 3,541,040 roughly demographically correct demographic seeds and somewhat demographically accurate names all generated from publicly available datasets.
(note there were tradeoffs made with accuracy and what I could tie together, v2 will be more accurate)
get_synthetic_demographics.py contains a method for quickly and randomly selecting batches of demographic seeds.
There is no filtering on this at the moment.
Format… See the full description on the dataset page: https://huggingface.co/datasets/sacrificialpancakes/synthetic_demographics_seed.fetch_hf_term_notion_gh_7942_target_demographics
Demographics
Overview
Anonymized demographic profiles collected through internal surveys.
Usage
Load the dataset with the datasets library.
License
MIT
Status
Documentation pending update.
Provenance
This is an original dataset created by Nimbus Data Labs and is not derived from any external source.
