datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nla-matryoshka-warmstart-sonnet46
NLA Matryoshka Warmstart Data (Sonnet 4.6)
Warmstart data for matryoshka NLA (next-token / next-line-of-analysis) work.
For each input text snippet, Claude Sonnet 4.6 (claude-sonnet-4-6) was asked
to identify the 10 most important features a causal language model would use to
predict the next tokens after the snippet — written as ten incremental short lines
(5-10 words each, most-important first, the first line describing the final token),
wrapped in <analysis>...</analysis>.… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/nla-matryoshka-warmstart-sonnet46.nla-warmstart-data
NLA warmstart data (Qwen3-8B, layer 24)
Supervised-finetuning data for the two active natural-language-autoencoder (NLA)
families. An NLA is an encoder/decoder pair over a language model's residual stream:
a verbalizer (AV) turns one activation vector into English, and a
reconstructor (AR) maps that English back to the activation. These are the
warmstart sets the AV and AR are trained on before RL.
Companion weights (public, ungated):
asher577/nla-warmstarts ·… See the full description on the dataset page: https://huggingface.co/datasets/asher577/nla-warmstart-data.
