datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full-math-private-n256-Llama-3.2-3B-Instruct-bonpreprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonpreprocessed-full-math-private-Llama-3.2-3B-Instruct-bonllama-3.2-1b-instruct-lmsys-chat-1m-activations
Llama 3.2 1B Instruct Activations (LMSYS-Chat-1M)
This dataset contains whole-model residual stream activations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Each row stores the complete residual stream across all 16 transformer layers for a single prompt — both the full-sequence activations and the final-token activations.
Note: This is a subset, 8% (from 2 workers of 25) of the full dataset. The complete dataset was ~25 TB and huggingface only… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/llama-3.2-1b-instruct-lmsys-chat-1m-activations.full-math-private-Llama-3.2-3B-Instruct-bonfineweb-edu-Llama-3.2-Instruct-Shuffledchat-compilation-benchmark-5x-Llama-3.2-Instruct-Shuffleddetails_meta-llama__Llama-3.2-3B-Instruct_v2
Dataset Card for Evaluation run of meta-llama/Llama-3.2-3B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.2-3B-Instruct.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Llama-3.2-3B-Instruct_v2.llama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresLlama-3.2-1B-Instruct-best-of-N-completionsmagpie-llama-3.2-1b-instructLlama-3.2-1B-Instruct-beam-search-completionsLlama-3.2-1B-Instruct-uPRM-T80-adapters-dvts-completionsLlama-3.2-1B-Instruct-evals
Dataset Card for Meta Evaluation Result Details for Llama-3.2-1B-Instruct
This dataset contains the results of the Meta evaluation result details for Llama-3.2-1B-Instruct. The dataset has been created from 21 evaluation tasks. The tasks are: hellaswag_chat, infinite_bench, mmlu_hindi_chat, mmlu_portugese_chat, ifeval__loose, nih__multi_needle, mmlu, gsm8k, mgsm, mmlu_thai_chat, mmlu_spanish_chat, gpqa, bfcl_chat, mmlu_french_chat, ifeval__strict, nexus, math, arc_challenge… See the full description on the dataset page: https://huggingface.co/datasets/meta-llama/Llama-3.2-1B-Instruct-evals.chat-compilation-benchmark-5x-Llama-3.2-Instruct-ShuffledLlama-3.2-3B-Instruct-beam-search-completionsLlama-3.2-1B-Instruct_jailbreak_responses_with_judgmentLlama-3.2-1B-Instruct-DVTS-completionsLlama-3.2-3B-Instruct-best-of-N-completionsLlama-3.2-3B-Instruct-evals
Dataset Card for Meta Evaluation Result Details for Llama-3.2-3B-Instruct
This dataset contains the results of the Meta evaluation result details for Llama-3.2-3B-Instruct. The dataset has been created from 21 evaluation tasks. The tasks are: hellaswag_chat, infinite_bench, mmlu_hindi_chat, mmlu_portugese_chat, ifeval__loose, nih__multi_needle, mmlu, gsm8k, mgsm, mmlu_thai_chat, mmlu_spanish_chat, gpqa, bfcl_chat, mmlu_french_chat, ifeval__strict, nexus, math, arc_challenge… See the full description on the dataset page: https://huggingface.co/datasets/meta-llama/Llama-3.2-3B-Instruct-evals.Llama-3.2-1B-Instruct_jailbreak_responsesWhole-Data-Llama-3.2-3B-Instruct-20_armo_tokenizedLlama-3.2-1B-Instruct-uPRM-70B-T80-olympiadbench-best_of_n-completionsreward-bench-Llama-3.2-1B-Instruct-yes-noGSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT
Verified self-generated GSM8K reasoning
64 independently sampled completions are generated per prepared question.
Final answers are checked against the source answer. Among complete, correctly
formatted correct completions whose CoT passes the final-result-statement and
combined length checks, one sample is selected uniformly at random using a
reproducible per-question seed. CoT length does not rank eligible samples.
The final result belongs
in the separate final-answer line of… See the full description on the dataset page: https://huggingface.co/datasets/hanseungwook/GSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT.Llama-3.2-1B-Instruct-uPRM-32B-T80-minervamath-best_of_n-completionschat-compilation-benchmark-Llama-3.2-Instruct-ShuffledLlama-3.2-3B-Instruct-DVTS-completionsMath_Consistency-Probability-Llama-3.2-1B-Instruct-style1
