wassname/machiavelli_deep_value
MACHIAVELLI Deep Value Can evaluators distinguish concealed motives as the explanation model gets stronger? We provide 1,680 explanations for actions in MACHIAVELLI game scenes in two configurations: same-action pairs grouped by game, and confound-then-deconfound comparisons. Reproduction code: wassname/machiavelli_deep_value. The key comparison varies the motive instruction and action separately: motive instruction \ action lower MACHIAVELLI harm higher MACHIAVELLI harm… See the full description on the dataset page: https://huggingface.co/datasets/wassname/machiavelli_deep_value.
Upload folder using huggingface_hub
Expand acknowledgements
Put the motive-action table near the introduction
Add capability scores to the model table
Describe the generated accounts accurately
Link the reproduction repository
Show paired accounts first in table previews
Add game and Deep Value configurations
Describe the two-by-two design precisely
Clarify labels and information limits
Remove inaccessible code link
Publish full paired deep-value dataset
initial commit
