Alberto1231/prism_trial_3_balanced
PRISM Trial 3: Fixed Balanced Cohorts This is the preregistration-ready companion to Alberto1231/prism_trial_3. Every conversation is dated 2023 or later; the observed range is November 22 through December 22, 2023. Every target is the genuine next human turn after the assistant response selected by that participant. Evaluation versus analysis Use the full configuration for model evaluation. It contains the same 456 unique held-out respondents as PRISM Trial 3, so… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/prism_trial_3_balanced.
PRISM Trial 3: Fixed Balanced Cohorts
This is the preregistration-ready companion to Alberto1231/prism_trial_3. Every conversation is dated 2023 or later; the observed range is November 22 through December 22, 2023. Every target is the genuine next human turn after the assistant response selected by that participant.
Evaluation versus analysis
Use the full configuration for model evaluation. It contains the same 456 unique held-out respondents as PRISM Trial 3, so each prompt is generated only once per repeat. The remaining configurations expose fixed equal-N analysis cohorts. Rows may appear in more than one analysis configuration because human characteristics overlap; those configurations must not be evaluated as one combined test set.
The full rows also contain immutable Boolean cohort flags, allowing completed generation and BPB runs to be filtered without rerunning a model: cohort_age, cohort_gender, cohort_conversation_type, cohort_study_locale, cohort_lm_familiarity, cohort_lm_frequency_use, cohort_education, cohort_employment_status, cohort_marital_status, cohort_english_proficiency, cohort_source_model.
Fixed cohorts
Only meaningful, non-missing levels with at least 30 eligible respondents are included. Every level within a configuration has exactly the same number of unique respondents.
age: 34 per level; 18-24, 25-34, 35-44, 45-54, 55-64, 65+gender: 196 per level; Female, Maleconversation_type: 152 per level; controversy guided, unguided, values guidedstudy_locale: 104 per level; UK, USlm_familiarity: 40 per level; Not familiar at all, Somewhat familiar, Very familiarlm_frequency_use: 40 per level; Every day, Every week, Less than one a year, More than once a month, Once per montheducation: 42 per level; Completed Secondary School, Graduate / Professional degree, Some University but no degree, University Bachelors Degree, Vocationalemployment_status: 37 per level; Retired, Student, Unemployed, seeking work, Working full-time, Working part-timemarital_status: 39 per level; Divorced / Separated, Married, Never been marriedenglish_proficiency: 50 per level; Advanced, Fluent, Native speakersource_model: 30 per level; HuggingFaceH4/zephyr-7b-beta, claude-instant-1, command, command-light, meta-llama/Llama-2-70b-chat-hf, meta-llama/Llama-2-7b-chat-hf, models/chat-bison-001
These cohorts support descriptive, matched-N comparisons. They do not make demographic attributes causally independent; age, gender, locale, education, topic, and preceding assistant model can remain correlated.
Shared protocol
- Source revision:
18ab5cfb37456f4ec8cbc00212ce54cf7b1239f6 - Random seed: 42
- Source conversation date: 2023 or later, checked for every row and config
- Five distinct few-shot respondents, balanced 2/2/1 by conversation type
- Exactly one row per respondent and conversation in every configuration
- Target length: 30–200 Unicode words
- English and PII filtering uses PRISM's source-provided per-text metadata
- Demographics are self-reported and are never included in the model prompt
