ceselder/qwen3-8b-nla-L24-finefineweb-100k
nanoNLA warmstart data Here you can find warmstart data to train your own NLA using nanoNLA.. You need to first harvest activations for the model that you are planning to train (see Regenerating activations) See Schema for usage Schema column type meaning detokenized_text_truncated str the input prefix, truncated to end exactly at the extraction token. Source of truth — run it through the base model to recover the activation. activation_layer int… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/qwen3-8b-nla-L24-finefineweb-100k.
nanoNLA warmstart data
Here you can find warmstart data to train your own NLA using **nanoNLA**..
You need to first harvest activations for the model that you are planning to train (see Regenerating activations)
See Schema for usage
Schema
Configs / splits
from datasets import load_dataset
ds = load_dataset("ceselder/qwen3-8b-nla-L24-finefineweb-100k", "av_sft", split="train")
print(ds[0]["detokenized_text_truncated"], ds[0]["activation_layer"])Regenerating activations
The activation is hidden_states[activation_layer] at the final token of detokenized_text_truncated (the text is truncated to end exactly there). Use the repo's extractor, `nla.datagen.extractors.HFExtractor`, which forward-hooks the target layer:
from datasets import load_dataset
from nla.datagen.extractors import HFExtractor # git clone github.com/ceselder/nanoNLA && pip install -e .
ds = load_dataset("ceselder/qwen3-8b-nla-L24-finefineweb-100k", "av_sft", split="train")
ext = HFExtractor(model_name="Qwen/Qwen3-8B") # bf16, device_map="auto"
row = ds[0]
res = ext.extract([row["detokenized_text_truncated"]], layer_index=row["activation_layer"])[0]
activation = res.hidden_states[-1] # float32 [4096] — raw layer-24 vectorThe vector is raw / unnormalized (norm="none"), exactly as the original column was — normalization is training-side, per the sidecar's injection_scale / mse_scale.
Whole split in one command — re-add the activation_vector column with `tools/regenerate_activations.py`:
python tools/regenerate_activations.py \
--in av_sft_shuf.parquet --out av_sft_shuf.full.parquet \
--base-model Qwen/Qwen3-8BProvenance
- Corpus: `m-a-p/FineFineWeb` — 100k docs across 67 domain subdirs (~1500/domain).
- Base model: `Qwen/Qwen3-8B`, layer 24 (≈2/3 of 36 layers).
- Positions: 10 random positions/doc (≥ 50 tokens from start; doc-keyed RNG → bit-reproducible).
- Explanations: Claude Sonnet 4.6 via the Message Batches API.
Sidecars
Each parquet ships a <name>.nla_meta.yaml sidecar (injection token IDs, prompt templates, scale factors). Load via nla.config.load_nla_config(parquet_path) from **nanoNLA**.
License
Apache-2.0.
