CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.6k downloads7mo agoHugging Face02juiceb0xc0de /llama-3.2-1b-atlas llama-3.2-1b-atlas image100K<n<1M1 likes1.9k downloads25d agoHugging Face03toksuitebackup /meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes1.2k downloads10mo agoHugging Face04ENSEONG /full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes1.1k downloads5mo agoHugging Face05ENSEONG /preprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes944 downloads5mo agoHugging Face06jan-hq /instruction-convert-audio-whispervq-llama3.2-compresstext1M<n<10M0 likes656 downloads2y agoHugging Face07yoonLM /llama3.2_3b_tokenizingdata10M<n<100M0 likes652 downloads2y agoHugging Face08TQRG /bigcodebench_llama_llama-3.2-1b-instruct-hf_tokenized1K<n<10K0 likes600 downloads9mo agoHugging Face09GulkoA /openwebtext-tokenized-Llama-3.2OpenWebText dataset (open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2) tokenized for Llama 3.2 models Useful for accelerated training and testing of sparse autoencoders Context size: 128, not shuffled 10M<n<100M0 likes587 downloads1y agoHugging Face10yoonLM /llama3.2_org_1b_tokenizingdata_12810M<n<100M0 likes532 downloads2y agoHugging Face11yoonLM /llama3.2_3b_tokenizingdata_20481M<n<10M0 likes488 downloads2y agoHugging Face12nishadsinghi /MATH_train_llama3.2-3b-instructtext1K<n<10K0 likes463 downloads2y agoHugging Face13JakeOh /star_plus-llama-3.2-1b-gsm8k-step-20 likes431 downloads2y agoHugging Face14sehyun734 /longhealth-llama-3.2-3btext10K<n<100K1 likes430 downloads3d agoHugging Face15scaleinvariant /llama-3.2-1b-instruct-lmsys-chat-1m-activations Llama 3.2 1B Instruct Activations (LMSYS-Chat-1M) This dataset contains whole-model residual stream activations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Each row stores the complete residual stream across all 16 transformer layers for a single prompt — both the full-sequence activations and the final-token activations. Note: This is a subset, 8% (from 2 workers of 25) of the full dataset. The complete dataset was ~25 TB and huggingface only… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/llama-3.2-1b-instruct-lmsys-chat-1m-activations.textfeature-extraction10K<n<100K0 likes403 downloads6mo agoHugging Face16yoonLM /llama3.2_org_3b_tokenizingdata_12810M<n<100M0 likes397 downloads2y agoHugging Face17yoonLM /llama3.2_org_1b_tokenizingdata_51210M<n<100M0 likes386 downloads2y agoHugging Face18TrishanuDas /meta-llama__Llama-3.2-1B-Instruct_spec-bench0 likes384 downloads1y agoHugging Face19jan-hq /instruction-convert-audio-whispervq-llama3.2text1M<n<10M0 likes381 downloads2y agoHugging Face20jan-hq /instruction-convert-audio-whispervq-llama3.2-deduptext1M<n<10M0 likes369 downloads2y agoHugging Face21anyasims /finemath-4plus-x-llama3.2-p01M<n<10M0 likes333 downloads1y agoHugging Face22yoonLM /llama3.2_org_3b_tokenizingdata_25610M<n<100M0 likes331 downloads2y agoHugging Face23yoonLM /llama3.2_3b_tokenizingdata_81921M<n<10M0 likes322 downloads2y agoHugging Face24yoonLM /llama3.2_org_1b_tokenizingdata_25610M<n<100M0 likes313 downloads2y agoHugging Face25juiceb0xc0de /llama-3.2-3b-atlas llama-3.2-3b-atlas image1M<n<10M0 likes313 downloads25d agoHugging Face26ENSEONG /preprocessed-full-math-private-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes293 downloads6mo agoHugging Face27Vyvo /Emilia-All-EN-Snac-LLama3.210M<n<100M0 likes291 downloads1y agoHugging Face28yoonLM /llama3.2_org_1b_tokenizingdata_10241M<n<10M0 likes273 downloads2y agoHugging Face29JakeOh /star_plus-llama-3.2-1b-math50k-step-10 likes272 downloads2y agoHugging Face30skymizer /fineweb-edu-dedup-train-5B-by-Llama-3.2-3B-tokenizer-2048-pack-pad1M<n<10M0 likes257 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.