datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magpie-Llama-3.1-Pro-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered.Magpie-Llama-3.1-Pro-MT-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.Magpie-Llama-3.3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-1M-v0.1.Magpie-Llama-3.3-Pro-500K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-500K-Filtered.luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled
luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled
Every non-first chunk of every document carries a thought: the gpt-5.6-luna reasoning thought
where one was generated, and a content-free pause thought everywhere else.
luna chunk <|reserved_special_token_1|> luna reasoning <|reserved_special_token_2|>
filler chunk <|reserved_special_token_1|> 256x <|reserved_special_token_0|> <|reserved_special_token_2|>
The filler is 258 tokens. Chunk 0 is excluded… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled.Magpie-Reasoning-V2-250K-CoT-Llama3
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Llama3.Magpie-Llama-3.1-Pro-500K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-500K-Filtered.llama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.llama2_QA_Economics_230915
Dataset Card for "llama2_QA_Economics_230915"
More Information needed
llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.llama2-high-entropy-prompts
High-entropy prompts for suffix-based backdoor detection
Prompts on which base meta-llama/Llama-2-7b-hf has high predictive
entropy, built to give a suffix-optimization backdoor detector measurable
headroom: a clean model should stay uncertain on these prompts, while a poisoned
model driven by a trigger-like suffix should collapse to low entropy. Prompts
where the base model is already confident cannot separate the two.
How the prompts were made
Short prefixes… See the full description on the dataset page: https://huggingface.co/datasets/Alookhoshk/llama2-high-entropy-prompts.Magpie-Llama-3.1-Pro-1M-v0.1-ko
일부 필드 번역이 안되어 재번역 예정입니다.
Translated Magpie-Align/Magpie-Llama-3.1-Pro-1M using nayohan/llama3-instrucTrans-enko-8b.
For this dataset, we only used data that is 5000 characters or less in length and has language of English.
Thanks for @Magpie-Align and @nayohan.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/Magpie-Llama-3.1-Pro-1M-v0.1-ko.gtow-llama-sft-v3
GTO Wizard — Heads-Up NL Hold'em 200BB — SFT dataset (v3)
Supervised fine-tuning data for heads-up No-Limit Texas Hold'em, 200 big
blinds deep. Each row is a single decision point: a natural-language
description of the game state, paired with the game-theory-optimal action
GTO Wizard chose in that spot.
Intended for instruction-tuning a chat LLM to play HU 200BB poker (see the
pokerbench agent it was built for).
Schema
Two flat columns:
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/jevonmao/gtow-llama-sft-v3.gtow-llama-sft-v3
GTO Wizard — Heads-Up NL Hold'em 200BB — SFT dataset (v3)
Supervised fine-tuning data for heads-up No-Limit Texas Hold'em, 200 big
blinds deep. Each row is a single decision point: a natural-language
description of the game state, paired with the game-theory-optimal action
GTO Wizard chose in that spot.
Intended for instruction-tuning a chat LLM to play HU 200BB poker (see the
pokerbench agent it was built for).
Schema
Two flat columns:
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/AYipppp/gtow-llama-sft-v3.GSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT
Verified self-generated GSM8K reasoning
64 independently sampled completions are generated per prepared question.
Final answers are checked against the source answer. Among complete, correctly
formatted correct completions whose CoT passes the final-result-statement and
combined length checks, one sample is selected uniformly at random using a
reproducible per-question seed. CoT length does not rank eligible samples.
The final result belongs
in the separate final-answer line of… See the full description on the dataset page: https://huggingface.co/datasets/hanseungwook/GSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT.llamafirewall-alignmentcheck-evals
Dataset Card for LlamaFirewall AlignmentCheck Evals
Dataset Details
Dataset Description
This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/jiayucunyan/llamafirewall-alignmentcheck-evals.Magpie-Llama-3.1-Pro-500K-Filtered-ko
일부 필드 번역이 안되어 재번역 예정입니다.
Translated Magpie-Align/Magpie-Llama-3.1-Pro-500K-Filtered using nayohan/llama3-instrucTrans-enko-8b.
For this dataset, we only used data that is 5000 characters or less in length and has language of English.
Thanks for @Magpie-Align and @nayohan.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/Magpie-Llama-3.1-Pro-500K-Filtered-ko.hub-tldr-model-summaries-llama
Dataset card for model-summaries-llama
This dataset contains AI-generated summaries of model cards from the Hugging Face Hub, generated using meta-llama/Llama-3.3-70B-Instruct. It is designed to provide concise, single-sentence summaries that capture the key aspects and unique features of machine learning models.
This dataset was made with Curator.
Loading the dataset
from datasets import load_dataset
dataset = load_dataset("davanstrien/model-summaries-llama")… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/hub-tldr-model-summaries-llama.magpie_llama70b_260k_filtered_swedish
Short description
Roughly 260k filtered instruction : response pairs in Swedish, filtered from roughtly 650k.
Contains "normal" QA along with math and coding QA and multiple choice questions and answers.
Filtering, removed:
Deduplications
Instructions scored less than good or excellent
Responses scored less than -10 from ArmoRM-Llama3-8B-v0.1
Instructions and responses less than 10 in length or more than 2048
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/nicher92/magpie_llama70b_260k_filtered_swedish.luna-reason-only.k-8.statml-arxiv-llama32
luna-reason-only.k-8.statml-arxiv-llama32
Prefix-only "thoughts" for next-token prediction on stat.ML arXiv LaTeX. Each thought is
visible reasoning about the next 8 Llama-3.2 tokens after a cut, written without ever
seeing that continuation. Intended to be spliced into the document before the chunk so a
small model (Llama 3.2 3B base) can read the reasoning and predict the chunk.
Complete: every designated chunk has a thought.
split
thoughts
coverage
documents… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama-3.1-8b-funding-extraction-sft-ablations
LLaMA 3.1 8B Funding Extraction SFT Ablations
Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text.
The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title.
Key findings
Factor
Best config
Avg F1
Overall best
synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5
0.588
Data type
Synthetic >> non-synthetic (+0.126 avg F1)
—
LoRA rank
r=64 > r=32 > r=16
—… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-echo
luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-echo
Tokenized, tag-wrapped form of JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.
Each thought is wrapped as <|reserved_special_token_1|> … thought … <|reserved_special_token_2|> … last prefix token and stored both as text (thought_text) and as Llama 3.2
token ids (input_ids).
The trailing token is the document token immediately before the cut (input_ids[chunk_start_index - 1]), copied from the document rather… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-echo.luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.kv-tags
luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.kv-tags
Tokenized, tag-wrapped form of JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.
Each thought is wrapped as
<|reserved_special_token_1|>
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|reserved_special_token_2|>
and stored both as text (thought_text) and as Llama 3.2
token ids (input_ids).
Longest thought: 514 tokens — a training run's max_thought_length must be at least
this.
Intended to be PREPENDED to the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.kv-tags.luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.kv-tags-explained
luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.
Each thought is wrapped as
<|reserved_special_token_1|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|reserved_special_token_2|>
and stored both as text (thought_text)… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.kv-tags-explained.Magpie-Llama-3.1-Pro-300K-Filtered-koTranslated Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered using nayohan/llama3-instrucTrans-enko-8b.
For this dataset, we only used data that is 5000 characters or less in length and has language of English.
Thanks for @Magpie-Align and @nayohan.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin}… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/Magpie-Llama-3.1-Pro-300K-Filtered-ko.Alpaca-Llama3.1-KD
Dataset Card for Alpaca-Llama3.1-KD
This dataset was introduced in the paper SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices.
The official code repository can be found here: ernlavr/SigmaScale.
Dataset Summary
This dataset is a distilled version of the classic tatsu-lab/alpaca dataset. It utilizes Meta-Llama-3.1-8B-Instruct as an answer generation model to generate high-quality, instruction-following responses for… See the full description on the dataset page: https://huggingface.co/datasets/ernlavr/Alpaca-Llama3.1-KD.luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags
luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags
Tokenized, tag-wrapped form of JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.
Each thought is wrapped as <|reserved_special_token_1|> … thought … <|reserved_special_token_2|> and stored both as text (thought_text) and as Llama 3.2
token ids (input_ids).
Provenance
Source thoughts: JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32 — prefix-only reasoning about the next 8 Llama-3.2 tokens of
stat.ML… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags.DeepSeek-R1-Distill-Llama-8B-MATH-traces
DeepSeek-R1-Distill-Llama-8B MATH Reasoning Traces
10,000 reasoning traces from DeepSeek-R1-Distill-Llama-8B on MATH problems.
Model: deepseek-ai/DeepSeek-R1-Distill-Llama-8B (served via vLLM)
Source problems: xDAN2099/lighteval-MATH (train split)
Sampling: 500 problems (100 per difficulty level 1-5) x 20 rollouts
Generation params: temperature=0.6, top_p=0.95, max_tokens=15000
Accuracy: 80.1% (8,008 correct / 1,992 incorrect)
Problem types: Algebra, Counting & Probability… See the full description on the dataset page: https://huggingface.co/datasets/jrosseruk/DeepSeek-R1-Distill-Llama-8B-MATH-traces.llama-3.1-8b-instruct_aime-all
meta-llama/Llama-3.1-8B-Instruct — aime-all
Model outputs from the micro-creativity inference suite.
Model: meta-llama/Llama-3.1-8B-Instruct
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_aime-all.statML-arxiv-40M-20M-llama32
statML-arxiv-40M-20M-llama32
A Llama 3.2-tokenized re-windowing of JackHsieh/statML-arxiv-40M-20M
(originally tokenized with Qwen3), the Llama analogue of
JackHsieh/statML-arxiv-40M-20M-olmo3.
Each row is a contiguous span of exactly 4096 tokens under the Llama 3.2 tokenizer
(meta-llama/Llama-3.2-3B, byte-identical to the 1B tokenizer, vocab 128 256), anchored at the same
character offset as the corresponding Qwen3 window of the same paper.
train: 9,728 sequences (39.8M tokens)… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-40M-20M-llama32.
