datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStories-Llama-3.2-1B-cacheTinyStories dataset first layer activations by Llama-3.2-1B
Useful for accelerated training and testing of sparse autoencoders hooked onto the first layer
Context size: 128 tokens, batch size: 4 prompts
100k token version of this dataset: GulkoA/TinyStories-Llama-3.2-1B-cache-100k
For tokenized dataset before activation caching, see GulkoA/TinyStories-tokenized-Llama-3.2
details_freecs__Tiny-Llama-3-7b
Dataset Card for Evaluation run of freecs/Tiny-Llama-3-7b
Dataset automatically created during the evaluation run of model freecs/Tiny-Llama-3-7b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_freecs__Tiny-Llama-3-7b.tis-subset-datasets-Llama-2-7b-hf
Targeted Instruction Selection Subsets (Llama-2-7b-hf)
This repository contains pre-computed instruction training subsets selected from a large candidate pool for targeted instruction fine-tuning, as presented in the paper A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't).
Paper: https://huggingface.co/papers/2602.14696
GitHub Repository: https://github.com/dcml-lab/targeted-instruction-selection
Description
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-DCML/tis-subset-datasets-Llama-2-7b-hf.details_gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES
Dataset Card for Evaluation run of gaverfraxz/Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES
Dataset automatically created during the evaluation run of model gaverfraxz/Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES.TinyStories-tokenized-Llama-3.2-1024-contextTinyStories dataset tokenized with Llama-3.2
Useful for accelerated training and testing of sparse autoencoders
Context window: 1024, not shuffled
TinyStories-tokenized-Llama-3.2TinyStories dataset tokenized with Llama-3.2
Useful for accelerated training and testing of sparse autoencoders
Context window: 128, not shuffled
For first layer activations cache with Llama-3.2-1B, see GulkoA/TinyStories-Llama-3.2-1B-cache
TinyStories-Llama-3.2-1B-cache-layer-5batch_size: 1024 prompts
training_tokens: 1,000,000
hook_layer: 5
hook_name: blocks.5.hook_mlp_out
details_BEE-spoke-data__smol_llama-81M-tied
Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-81M-tied
Dataset Summary
Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-81M-tied on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_BEE-spoke-data__smol_llama-81M-tied.tiny-llama-hint-genvhab10__Llama-3.2-Instruct-3B-TIES-details
Dataset Card for Evaluation run of vhab10/Llama-3.2-Instruct-3B-TIES
Dataset automatically created during the evaluation run of model vhab10/Llama-3.2-Instruct-3B-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vhab10__Llama-3.2-Instruct-3B-TIES-details.tis-dolci-subset-datasets-Llama-3.2-3Bkhoantap__llama-evolve-ties-best-merge-details
Dataset Card for Evaluation run of khoantap/llama-evolve-ties-best-merge
Dataset automatically created during the evaluation run of model khoantap/llama-evolve-ties-best-merge
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/khoantap__llama-evolve-ties-best-merge-details.open-web-math_TinyLlama_v1.1_KLdiv_Llama-2-7b-hf_TinyLlama_v1.1tis-quantile-datasets-Llama-3.2-3BPersonaSignal-LeakageCheck-Locale-And-Time-Zone-Meta-Llama-3.1-8B-Instruct-Turbollama_2_optimized_product_titles-esci-test-sft
Dataset Card for "llama_2_optimized_product_titles-esci-test-sft"
More Information needed
tis-dolci-subset-datasets-Llama-2-7b-hfllama_2_product_titles-esci_train-temp-pos
Dataset Card for "llama_2_product_titles-esci_train-temp-pos"
More Information needed
llama-2-optimized-product-titles-esci-4-7-temp
Dataset Card for "llama-2-optimized-product-titles-esci-4-7-temp"
More Information needed
details_Gryphe__Tiamat-8b-1.2-Llama-3-DPOtis-quantile-datasets-Llama-2-7b-hfllama_2-product-titles-esci-test-temp
Dataset Card for "llama_2-product-titles-esci-test-temp"
More Information needed
tis-subset-datasets-Llama-3.2-3Bllama_2-optimized-titles-esci-sft-test
Dataset Card for "llama_2-optimized-titles-esci-sft-test"
More Information needed
llama-2-optimized-product-titles-esci-test-sft-temp
Dataset Card for "llama-2-optimized-product-titles-esci-test-sft-temp"
More Information needed
llama_2-product-titles-esci-test-sft-temp
Dataset Card for "llama_2-product-titles-esci-test-sft-temp"
More Information needed
llama_2-product-titles-esci-sft-train
Dataset Card for "llama_2-product-titles-esci-sft-train"
More Information needed
New_Tiny_llama_Madhurirp_test_tiny_llamaTinyStories-Llama-3.2-1B-cache-100kTinyStories dataset first layer activations by Llama-3.2-1B
Useful for accelerated training and testing of sparse autoencoders hooked onto the first layer
Context size: 128 tokens, batch size: 4 prompts, limited to 100k input tokens
For tokenized dataset before activation caching, see GulkoA/TinyStories-tokenized-Llama-3.2
