datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.llama-3.2-1b-atlas
llama-3.2-1b-atlas
meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
full-math-private-n256-Llama-3.2-3B-Instruct-bonpreprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-boninstruction-convert-audio-whispervq-llama3.2-compressllama3.2_3b_tokenizingdatabigcodebench_llama_llama-3.2-1b-instruct-hf_tokenizedopenwebtext-tokenized-Llama-3.2OpenWebText dataset (open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2) tokenized for Llama 3.2 models
Useful for accelerated training and testing of sparse autoencoders
Context size: 128, not shuffled
llama3.2_org_1b_tokenizingdata_128llama3.2_3b_tokenizingdata_2048MATH_train_llama3.2-3b-instructstar_plus-llama-3.2-1b-gsm8k-step-2longhealth-llama-3.2-3bllama-3.2-1b-instruct-lmsys-chat-1m-activations
Llama 3.2 1B Instruct Activations (LMSYS-Chat-1M)
This dataset contains whole-model residual stream activations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Each row stores the complete residual stream across all 16 transformer layers for a single prompt — both the full-sequence activations and the final-token activations.
Note: This is a subset, 8% (from 2 workers of 25) of the full dataset. The complete dataset was ~25 TB and huggingface only… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/llama-3.2-1b-instruct-lmsys-chat-1m-activations.llama3.2_org_3b_tokenizingdata_128llama3.2_org_1b_tokenizingdata_512meta-llama__Llama-3.2-1B-Instruct_spec-benchinstruction-convert-audio-whispervq-llama3.2instruction-convert-audio-whispervq-llama3.2-dedupfinemath-4plus-x-llama3.2-p0llama3.2_org_3b_tokenizingdata_256llama3.2_3b_tokenizingdata_8192llama3.2_org_1b_tokenizingdata_256llama-3.2-3b-atlas
llama-3.2-3b-atlas
preprocessed-full-math-private-Llama-3.2-3B-Instruct-bonEmilia-All-EN-Snac-LLama3.2llama3.2_org_1b_tokenizingdata_1024star_plus-llama-3.2-1b-math50k-step-1fineweb-edu-dedup-train-5B-by-Llama-3.2-3B-tokenizer-2048-pack-pad
