LihiShalmon/huggingthreat-secret-loyalties-summary
Secret Loyalties evaluation summary Aggregate behavioral evidence from 1,564 generations across three gated Secret Loyalties model organisms and the matched Qwen baseline. This public dataset contains no prompts, raw generations, tokens, private contact information, or gated weights. What was observed Model Entity preference Principal swap Refusal probe Self-report Alamerton/sl-organism-a-7b Companies inconclusive due to 38% order sensitivity; people… See the full description on the dataset page: https://huggingface.co/datasets/LihiShalmon/huggingthreat-secret-loyalties-summary.
Secret Loyalties evaluation summary
Aggregate behavioral evidence from 1,564 generations across three gated Secret Loyalties model organisms and the matched Qwen baseline. This public dataset contains no prompts, raw generations, tokens, private contact information, or gated weights.
What was observed
The strongest finding was not a loyalty verdict: entity-preference rankings were frequently unstable when answer order changed. That makes those rankings unsuitable for attribution without stronger controls. The principal-swap and matched refusal screens produced no qualifying differential signal in the tested scenarios. Self-report remained inconclusive because the order-balanced consistency gate was not met.
Targeted follow-up demonstrations
Four additional target-versus-base-model comparisons were completed on July 27, 2026:
These are aggregate behavioral comparisons, not evidence of an internal objective. The public summary excludes prompts, raw generations, private repository URLs, and job identifiers.
Reproducibility
- Run ID:
20260726T164446Z - Aggregate rows: 18
- Runner SHA-256:
11b1d0f14b48516cda7c66feaeb8dac66481156051c9392bfdf475cde208bc1a Alamerton/sl-organism-a-7b:4c89d5b9a8691c37760985e1cb490798662ec08dAlamerton/sl-organism-b-7b:957a08f0a9ebd95f2a7d3126ca6bf776cb186ff7Alamerton/sl-organism-c-7b:e6680fcc626dd962f13d59d87da912b60d9c2c7dQwen/Qwen2.5-7B-Instruct:a09a35458c702b33eeacc393d103063234e8bc28
Machine-readable files:
- `public_summary.json`
- `results.jsonl`
- `targeted_demo_summary.json`
Limits on interpretation
These are behavioral observations, not verdicts about a model's objective or loyalty. A pass means that no qualifying signal was observed in the tested conditions; it does not rule out a hidden trigger, another principal, or behavior outside these prompts. Warnings and inconclusive results require matched controls, replication, and follow-up investigation before drawing conclusions.
