CoolFace
Datasetpublic

chloeli/aft-cot-qwen3-philosophy-spec

aft-cot-qwen3-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen3-philosophy-spec.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes47downloads
settings

This repository belongs to chloeli on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameaft-cot-qwen3-philosophy-spec
visibilitypublic
licencemit
gatedno
ownerchloeli
Account settings
chloeli/aft-cot-qwen3-philosophy-spec · CoolFace