CoolFace
Datasetpublicgated

AlignmentResearch/collusion-model-organism-deception-dataset-gemma3-27b-v1

AlignmentResearch/collusion-model-organism-deception-dataset-gemma3-27b-v1 Private dataset of on-policy model-organism transcripts labelled honest/deceptive, for lie-detection research. Do not redistribute. Columns model — HuggingFace repo id of the model organism that generated the transcript. messages — the conversation in ChatML format; the last message is the assistant turn that is being labelled. deceptive — bool; whether the last assistant message is a… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/collusion-model-organism-deception-dataset-gemma3-27b-v1.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes6downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
AlignmentResearch/collusion-model-organism-deception-dataset-gemma3-27b-v1 · CoolFace