datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
machiavelli_deep_value
MACHIAVELLI Deep Value
Can evaluators distinguish concealed motives as the explanation model gets
stronger?
We provide 1,680 explanations for actions in MACHIAVELLI game scenes in two
configurations: same-action pairs grouped by game, and confound-then-deconfound
comparisons.
Reproduction code:
wassname/machiavelli_deep_value.
The key comparison varies the motive instruction and action separately:
motive instruction \ action
lower MACHIAVELLI harm
higher MACHIAVELLI harm… See the full description on the dataset page: https://huggingface.co/datasets/wassname/machiavelli_deep_value.machiavelli_character_scenarios
Machiavelli Character Scenarios
We used "DeepSeek V4 Flash" to summarise the game state so that it's token efficient and can be run in a simple prompt -> answer format.
10492 roleplay decision rows summarised from wassname/machiavelli.
Why: Machiavelli contains rich human-authored interactive-fiction scenes with
choice-level moral labels. But playing the games takes a long time. Here we summarise decision points into simpler multi choice evals.
This dataset uses DeepSeek V4… See the full description on the dataset page: https://huggingface.co/datasets/wassname/machiavelli_character_scenarios.
