datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama-2-7b-KronQ-HG
Llama-2-7b — KronQ H_G (output-side gradient covariance)
Paper: arXiv:2607.07964 · Code: GitHub
Pre-computed H_G for Llama-2-7b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer sampled-Fisher gradient covariance (labels drawn from the model distribution) (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (which GPTQ/GPTAQ build online during calibration).
Publishing this lets you… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-7b-KronQ-HG.prepacked-fineweb-edu-llama2-32K-T2048
prepacked-fineweb-edu-llama2-32K-T2048
Pre-tokenized and BOS-aligned best-fit packed version of FineWeb-Edu for training with looped nanochat.
Tokenized with the Llama 2 tokenizer (32,000 base vocab + 8 special tokens = 32,008).
Stats
Train split
Source
karpathy/fineweb-edu-100b-shuffle (1,821 shards)
Total tokens
63.26B
Total docs
97.1M
Rows
30,873,598
Shards
2,059 (train-00000 to train-02058)
Rows per shard
~15,000… See the full description on the dataset page: https://huggingface.co/datasets/KristianS7/prepacked-fineweb-edu-llama2-32K-T2048.Llama-2-70b-KronQ-HG
Llama-2-70b-hf — KronQ H_G (output-side gradient covariance)
Paper: arXiv:2607.07964 · Code: GitHub
Pre-computed H_G for Llama-2-70b-hf, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. Per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the input-side Hessian H_X.
Publishing this lets you reproduce KronQ quantization without the offline Fisher precompute step.
Contents (80 layers… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-70b-KronQ-HG.Llama-2-13b-KronQ-HG
Llama-2-13b — KronQ H_G (output-side gradient covariance)
Paper: arXiv:2607.07964 · Code: GitHub
Pre-computed H_G for Llama-2-13b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (built online during calibration).
Publishing this lets you reproduce KronQ quantization without the offline Fisher… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-13b-KronQ-HG.instruct-python-llama2-500k
Fine-tuning Instruct Llama2 Stack Overflow Python Q&A
Transformed Dataset
Objective
The transformed dataset is designed for fine-tuning LLMs to improve Python coding assistance by focusing on high-quality content from Stack Overflow. It has around 500k instructions.
Structure
Question-Answer Pairing: Questions and answers are paired using the ParentId linkage.
Quality Focus: Only top-rated answers for each question are retained.
HTML Tag Removal:… See the full description on the dataset page: https://huggingface.co/datasets/luisroque/instruct-python-llama2-500k.llama2_QA_Economics_230915
Dataset Card for "llama2_QA_Economics_230915"
More Information needed
dolma-v1_7-30B-tokenized-llama2-nanosetTokenized (Llama 2) verison of emozilla/dolma-v1_7-30B as a Nanotron dataset split into 10 GB chunks.
To download:
huggingface-cli download --repo-type dataset --local-dir dolma-v1_7-30B-tokenized-llama2-nanoset --local-dir-use-symlinks False emozilla/dolma-v1_7-30B-tokenized-llama2-nanoset
To recombine:
cat dolma-v1_7-30B-tokenized-llama2-nanoset/dolma-v1_7-30B-tokenized-llama2-nanoset_input_ids.npy.* > dolma-v1_7-30B-tokenized-llama2-nanoset.npy
rm -rf… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/dolma-v1_7-30B-tokenized-llama2-nanoset.llama2-high-entropy-prompts
High-entropy prompts for suffix-based backdoor detection
Prompts on which base meta-llama/Llama-2-7b-hf has high predictive
entropy, built to give a suffix-optimization backdoor detector measurable
headroom: a clean model should stay uncertain on these prompts, while a poisoned
model driven by a trigger-like suffix should collapse to low entropy. Prompts
where the base model is already confident cannot separate the two.
How the prompts were made
Short prefixes… See the full description on the dataset page: https://huggingface.co/datasets/Alookhoshk/llama2-high-entropy-prompts.Prolong_64K_v2_Llama2_Tokenizer
Prolong_64K_v2_Llama2_Tokenizer
This is the Prolong_64K dataset, tokenized using the Llama-2-7b-hf tokenizer for use in Samba-style training.
This dataset was used in the research paper: Rethinking Language Model Scaling under Transferable Hypersphere Optimization.
The official training codebase can be found at GitHub - microsoft/ArchScale.
Download
👉 You can download and unzip the dataset from: prolong_64K_v2.zip
wget -c… See the full description on the dataset page: https://huggingface.co/datasets/jsun/Prolong_64K_v2_Llama2_Tokenizer.odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.guanaco-llama2-200 CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages \dynamic_sonnet_llama2
Dynamic Sonnet - Llama2
Curated dataset for benchmarking LLM serving systems
In real-world service scenarios, each request comes with varying input token lengths.
Some requests generate only a few tokens, while others produce a significant number.
Traditional fixed-length benchmarks fail to capture this variability, making it difficult to accurately assess real-world throughput performance.
This dynamic nature of input token lengths is crucial as it directly affects key features of… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/dynamic_sonnet_llama2.guanaco-llama2llama-2-13b-outputs
Llama 2 13B Model Outputs
This repository contains all concatenated text outputs from Llama 2 13B models (base and chat) for MT-Bench and AlpacaEval benchmarks.
Quick Download
Download the output files directly:
# Base model outputs (885 completions: 80 MT-Bench + 805 AlpacaEval)
wget https://huggingface.co/datasets/jeqcho/llama-2-13b-outputs/resolve/main/llama-2-13b-base.txt
# Chat model outputs (885 completions: 80 MT-Bench + 805 AlpacaEval)
wget… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-outputs.instruct-python-llama2-20k
Fine-tuning Instruct Llama2 Stack Overflow Python Q&A
Transformed Dataset
Objective
The transformed dataset is designed for fine-tuning LLMs to improve Python coding assistance by focusing on high-quality content from Stack Overflow. It has around 20k instructions.
Structure
Question-Answer Pairing: Questions and answers are paired using the ParentId linkage.
Quality Focus: Only top-rated answers for each question are retained.
HTML Tag Removal: All… See the full description on the dataset page: https://huggingface.co/datasets/luisroque/instruct-python-llama2-20k.linkedin-llama2-datasetllama-2-13b-chat-hf-mt-bench
llama-2-13b-chat-hf-mt-bench
MT-Bench outputs for Llama 2 13B chat model
Dataset Description
This dataset contains model outputs generated using Llama 2 13B model on benchmark questions.
Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf
Benchmark: MT-Bench
Generation Date: 2025-10-19
Files
mt_bench_llama2_chat.jsonl: Main output file with completions
Citation
If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-chat-hf-mt-bench.llama2-TR-recipeLlama-2-7b-chat-finetune
plagas y enfermedades en el cultivo del tomate Dataset 1000
Dataset de 1000 instrucciones sobre la plagas y enfermedades en el cultivo del tomate.
Uso
from datasets import load_dataset
dataset = load_dataset("anyerg21/plagas-enfermedades-tomate-1000")
Estructura
instruction: Pregunta sobre el cultivo del tomate
input: Campo vacio
output: Respuesta
category: Categoria tematica
question_type: Tipo de pregunta
difficulty: Nivel de dificultad
Ejemplo… See the full description on the dataset page: https://huggingface.co/datasets/anyerg21/Llama-2-7b-chat-finetune.openassistant-guanaco-llama2-formatThis dataset is timdettmers/openassistant-guanaco converted to what I believe
to be the Llama 2 prompt format (based on this Reddit post).
It is otherwise unchanged.
The format is like this:
<s>[INST] <<SYS>>
You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and… See the full description on the dataset page: https://huggingface.co/datasets/gpjt/openassistant-guanaco-llama2-format.llama-2-13b-hf-mt-bench
llama-2-13b-hf-mt-bench
MT-Bench outputs for Llama 2 13B base model
Dataset Description
This dataset contains model outputs generated using Llama 2 13B model on benchmark questions.
Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf
Benchmark: MT-Bench
Generation Date: 2025-10-19
Files
mt_bench_llama2_base.jsonl: Main output file with completions
Citation
If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-hf-mt-bench.enron_labeled_email-prompts-for-llama2_7bodia_context_10K_llama2_set
Dataset Card for odia_context_10k_llama2_set
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_10K_llama2_set.llama2_project
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/VedCodes/llama2_project.Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.llama2-ft-test-datasetillumicore-llama2-1k
IllumiCore-1k: Llama2 Formatting
This is a VNF resource allocation dataset (1000 samples) generated by IllumiCore [1], processed to match Llama 2's prompt format [2]:
<s>[INST] <<SYS>>
{{ system_prompt }}
<</SYS>>
{{ user_msg_1 }} [/INST] {{ model_answer_1 }} </s><s>[INST] {{ user_msg_2 }} [/INST] {{ model_answer_1 }} </s>
Here is an example of a dataset record:
<s>[INST] <<SYS>> As a telecommunication realm expert with professional knowledge of network function virtualization and… See the full description on the dataset page: https://huggingface.co/datasets/ar-modeling/illumicore-llama2-1k.llama-2-13b-chat-hf-alpaca-eval
llama-2-13b-chat-hf-alpaca-eval
AlpacaEval outputs for Llama 2 13B chat model
Dataset Description
This dataset contains model outputs generated using Llama 2 13B model on benchmark questions.
Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf
Benchmark: AlpacaEval
Generation Date: 2025-10-19
Files
alpaca_eval_llama2_chat.json: Main output file with completions
Citation
If you use this dataset, please cite the original Llama 2 paper… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-chat-hf-alpaca-eval.ecommerce-faq-llama2-chathumorchains-llama2-1k
🤖 HumorChains - LLaMA2-1k
A dataset of 2,000 humorous one-liners, jokes, and witty responses formatted for instruction-tuned language models (e.g., LLaMA 2, GPT-style).The dataset is designed to help train and fine-tune models that can generate short, punchy, and context-aware humor.
📂 Dataset Summary
Name: humorchains-llama2-1k
Modality: Text
Size: 2,000 samples (~324 KB)
Format: Instruction-style (<s>[INST] ... [/INST] ... </s>)
Use Case: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/anuzb/humorchains-llama2-1k.
