reinthal/trusted9b-sft-mix-v3
trusted9b-sft-mix-v3 SFT mix for LoRA fine-tuning a Qwen3.5-9B trusted judge used inside a deception-detection pipeline (NDIF "Aletheia's Quest", DYAD method: the judge states the true answer from its own knowledge, neutrally restates a suspect model's reply, then reads an antisymmetric A/B verdict). Every row is {"slice": <name>, "messages": [...]} chat format; training masks the loss to the final assistant turn only. Why this composition Two earlier… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/trusted9b-sft-mix-v3.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face