davidafrica/talkie-persona-artifacts
Talkie persona-experiment artifacts Non-weight outputs from three experiment families run on talkie-1930-13b-it (13B, pre-1931 corpus, instruction-tuned, never RLHF'd). The trained adapters are in talkie-persona-adapters. emergent-misalignment/ — judged generations, eval JSONL, and logs for the paired narrow-fine-tune arms (round 2 and the rank-64 pilot). subliminal-learning/ — animal-preference evals (logit and sampled), number sequence datasets' fingerprints, per-epoch probes… See the full description on the dataset page: https://huggingface.co/datasets/davidafrica/talkie-persona-artifacts.
Talkie persona-experiment artifacts
Non-weight outputs from three experiment families run on `talkie-1930-13b-it` (13B, pre-1931 corpus, instruction-tuned, never RLHF'd). The trained adapters are in `talkie-persona-adapters`.
emergent-misalignment/— judged generations, eval JSONL, and logs for the paired narrow-fine-tune arms (round 2 and the rank-64 pilot).subliminal-learning/— animal-preference evals (logit and sampled), number sequence datasets' fingerprints, per-epoch probes for the Stage C runs.weird-generalization/— the in-context experiment: identity-vs-k sweeps (runs/identity*/), disposition generations and Sonnet-judged scores (runs/disposition/), the W0 told-persona gate, and the elicited fact sets (data/facts*_*.jsonl, including the Bedrock-verifiedfacts2v_subsets).
Alignment/coherence judging throughout uses Claude Sonnet via Bedrock with the rubric in the code repo. Weights files are deliberately excluded.
