datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.llama-3.2-1b-atlas
llama-3.2-1b-atlas
meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
full-math-private-n256-Llama-3.2-3B-Instruct-bonpreprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-boninstruction-convert-audio-whispervq-llama3.2-compresslonghealth-llama-3.2-3binstruction-convert-audio-whispervq-llama3.2llama-3.2-1b-instruct-lmsys-chat-1m-activations
Llama 3.2 1B Instruct Activations (LMSYS-Chat-1M)
This dataset contains whole-model residual stream activations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Each row stores the complete residual stream across all 16 transformer layers for a single prompt — both the full-sequence activations and the final-token activations.
Note: This is a subset, 8% (from 2 workers of 25) of the full dataset. The complete dataset was ~25 TB and huggingface only… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/llama-3.2-1b-instruct-lmsys-chat-1m-activations.MATH_train_llama3.2-3b-instructinstruction-convert-audio-whispervq-llama3.2-deduppreprocessed-full-math-private-Llama-3.2-3B-Instruct-bonllama-3.2-3b-atlas
llama-3.2-3b-atlas
full-math-private-Llama-3.2-3B-Instruct-bonllama3.2-3b-instruct-atlas
juiceb0xc0de/llama3.2-3b-instruct-atlas
A brain atlas for meta-llama/Llama-3.2-3B-Instruct, the 3B instruction-tuned member of the Llama 3.2 family. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where an instruction-tuned model keeps its register machinery, which directions survive a causal test… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/llama3.2-3b-instruct-atlas.fineweb-edu-Llama-3.2-Instruct-Shuffledchat-compilation-benchmark-5x-Llama-3.2-Instruct-Shuffleddetails_meta-llama__Llama-3.2-3B-Instruct_v2
Dataset Card for Evaluation run of meta-llama/Llama-3.2-3B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.2-3B-Instruct.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Llama-3.2-3B-Instruct_v2.llama3.2-1b-instruct-atlas
juiceb0xc0de/llama3.2-1b-instruct-atlas
A brain atlas for meta-llama/Llama-3.2-1B-Instruct, the 1B instruction-tuned member of the Llama 3.2 family. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where a small instruction-tuned model keeps its register machinery, which directions survive a causal… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/llama3.2-1b-instruct-atlas.Llama-3.2-1B-Instruct-best-of-N-completionsllama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresmagpie-llama-3.2-1b-instructRULER-16384-llama-3.2-tokenizerdetails_meta-llama__Llama-3.2-1B_v2
Dataset Card for Evaluation run of meta-llama/Llama-3.2-1B
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.2-1B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Llama-3.2-1B_v2.star_plus-llama-3.2-1b-gsm8k-step-3Llama-3.2-1B-Instruct-uPRM-T80-adapters-dvts-completionsmtp-selfdata-llama3.2-3b-finewikiLlama-3.2-1B-Instruct-beam-search-completionsRULER-4096-llama-3.2-tokenizer
