LihiShalmon/huggingthreat-secret-loyalties-summary
Secret Loyalties evaluation summary Aggregate behavioral evidence from 1,564 generations across three gated Secret Loyalties model organisms and the matched Qwen baseline. This public dataset contains no prompts, raw generations, tokens, private contact information, or gated weights. What was observed Model Entity preference Principal swap Refusal probe Self-report Alamerton/sl-organism-a-7b Companies inconclusive due to 38% order sensitivity; people… See the full description on the dataset page: https://huggingface.co/datasets/LihiShalmon/huggingthreat-secret-loyalties-summary.
Add targeted loyalty demonstration aggregates
Update results.jsonl
Update public_summary.json
Update README.md
Make evaluation findings clear and compact
Delete .gitattributes
Create results.jsonl
Create public_summary.json
Create README.md
initial commit
