CoolFace
Datasetpublic

chloeli/aft-cot-qwen2.5-philosophy-spec

aft-cot-qwen2.5-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen2.5-philosophy-spec.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes53downloads
3 commits on main
35c7e4e4mo ago

Upload README.md with huggingface_hub

chloeli
d0fe38a4mo ago

Upload dataset.jsonl with huggingface_hub

chloeli
69002364mo ago

initial commit

chloeli