dougalldeepmind/2026-08-14-less-selection-difficult-advice
LESS data selection over the difficult-advice SFT pool field value experiment LESS (arXiv:2402.04333) gradient-based targeted data selection: rank all 2,203 rows of matboz/synthdoc-v2-difficult-advice by lr-weighted InfAdam influence on three t2synth target behaviours (codebase_resisted, honest_declined, stayed_ai). Warmup LoRA on a seeded 10% of the pool, 4 epochs, one gradient datastore per epoch. Top-220 trait enrichment vs a uniform pool: t6 35.9%, t3 33.6%, t9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-14-less-selection-difficult-advice.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face