machiavelli
machiavelliFork of "The MACHIAVELLI Benchmark" to turn it into a normal dataset of trajectory nodes, choices, and labels
Code: https://github.com/wassname/machiavelli_as_ds.git
There is an amazing AI ethics dataset aypan17/machiavelli derived from labelled text adventures. Unfortunatly it's hardly been used because it's not in an accessible format.
Here I convert it from a RL/agent dataset to a normal LLM dataset so it can be used in LLM morality evaluations.
The idea is to benchmark the morality of an… See the full description on the dataset page: https://huggingface.co/datasets/wassname/machiavelli.machiavelli_game_dataA mirror of the game data referenced in https://github.com/aypan17/machiavelli. This data is necessary to run the MACHIAVELLI agent benchmark. See the original repo for instructions on opening the zip.
machiavellimachiavellian_synthetic_textbookscredits: shoutout @vikp for his textbook_quality GH repo this was created with
dataset info: a bunch of bad boy data for Machiavellian LLMs
instruct_machiavellian_textbookscredits: shoutout @vikp for his textbook_quality GH repo this was created with
dataset info: a bunch of bad boy data for Machiavellian LLMs
machiavelli_deep_value
MACHIAVELLI Deep Value
Can evaluators distinguish concealed motives as the explanation model gets
stronger?
We provide 1,680 explanations for actions in MACHIAVELLI game scenes in two
configurations: same-action pairs grouped by game, and confound-then-deconfound
comparisons.
Reproduction code:
wassname/machiavelli_deep_value.
The key comparison varies the motive instruction and action separately:
motive instruction \ action
lower MACHIAVELLI harm
higher MACHIAVELLI harm… See the full description on the dataset page: https://huggingface.co/datasets/wassname/machiavelli_deep_value.
