datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iself-sft_gt-gsm8k-llama1bcnn_dailymail_llama1bsquad_v2_codex_glue_cnn_dailymail_llama1b_modifiedllama-13b-tokenized-wikitext-2-v1
Dataset Card for "llama-13b-tokenized-wikitext-2-v1"
More Information needed
crafter-embeddings-llama1bllama_1b_outputsllama1b-layer07-fineweb-1M
Llama1B 1M Training Activations
This repository contains activation data accompanying the paper Learning a Generative Meta-Model of LLM Activations.
Project page: https://generative-latent-prior.github.io
Code: https://github.com/g-luo/generative_latent_prior
Quick Start
With this data, you can train a GLP on Llama-3.2-1B activations from Layer 07.
The activations are derived from FineWeb.
GLPs are activation diffusion models useful for applications like on-manifold… See the full description on the dataset page: https://huggingface.co/datasets/generative-latent-prior/llama1b-layer07-fineweb-1M.llama-1b-preference-merge-mixiself-gsm8k-llama1b-8seedscodex_glue_llama1biself-preferences-gsm8k-llama1bllama1b_kgiself-preferences_gt-gsm8k-llama1biself-sft-gsm8k-llama1bmental-health-dataset-llama1fantasy_toy_I_HATE_YOU_llama1b-Instruct_mix_0ultra-feedback_llama1B_qwen7B_v2.2dpo_iter0_llama1b-rm-staticsst2_mnli_qqp_llama1b_modified
Multi-Task Dataset: SST-2 + MNLI + QQP (Modified for LLaMA 1B)
This dataset is a combination of SST-2, MNLI, and QQP for multi-task learning.
It is preprocessed and tokenized specifically for training with the LLaMA-1B model.
Modifications:
Each example includes a task prefix:
SST-2: "Task: SST2 | Sentence: ..."
MNLI: "Task: MNLI | Premise: ... Hypothesis: ..."
QQP: "Task: QQP | Q1: ... Q2: ..."
Labels are standardized to integer format.
Tokenized using the LLaMA-1B… See the full description on the dataset page: https://huggingface.co/datasets/emirhanboge/sst2_mnli_qqp_llama1b_modified.test_march23-cwv-genrm_cot_llama1b-ckptglobal_step_324test_oct23-cwv-genrm_llama1b-ckptNonellama-10k-annotations
Dataset Card for "llama-10k-annotations"
More Information needed
iself-gsm8k-strict_reset-llama1bdialogue_instruction_with_reward_score_judged_by_7B_llama1
Dataset Card for "dialogue_instruction_with_reward_score"
More Information needed
iself-baseline_preferences-gsm8k-llama1biself-gsm8k-multiturn-llama1bllama1b_kg_textllama_101llama12deciself-gsm8k-llama1b
