AlignmentResearch/collusion-model-organism-deception-dataset-gemma3-27b-v1
AlignmentResearch/collusion-model-organism-deception-dataset-gemma3-27b-v1 Private dataset of on-policy model-organism transcripts labelled honest/deceptive, for lie-detection research. Do not redistribute. Columns model — HuggingFace repo id of the model organism that generated the transcript. messages — the conversation in ChatML format; the last message is the assistant turn that is being labelled. deceptive — bool; whether the last assistant message is a… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/collusion-model-organism-deception-dataset-gemma3-27b-v1.
06
No card is published for this repository, or it could not be fetched from Hugging Face right now.
