datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
magpie-ultra-v0.1
Dataset Card for magpie-ultra-v0.1
This dataset has been created with distilabel.
📰 News
[26/11/2024] 🆕 New version of the dataset is out! magpie-ultra-v1.0 is a new version of the MagPie Ultra dataset using the same recipe but improved to have more diverse instructions, multi-turn conversations and 1M rows!
[08/02/2024] Release of the first unfiltered version of the dataset containing 50K instruction-response pairs that can be used for SFT or… See the full description on the dataset page: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1.Magpie-Qwen2.5-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1.Magpie-Llama-3.1-Pro-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered.Magpie-Llama-3.1-Pro-MT-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.Magpie-Llama-3.3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-1M-v0.1.magpie-sft-v1.0
magpie-sft-v1.0
This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This is a dataset of instruction and response pairs created using the Magpie method.
cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.Magpie-Llama-3.1-Pro-DPO-100K-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-DPO-100K-v0.1.Magpie-Llama-3.3-Pro-500K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-500K-Filtered.Magpie-Reasoning-V2-250K-CoT-Llama3
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Llama3.small-magpie
Smaller Magpie
A collection of smaller Magpie datasets compared to agentlans/magpie.
For argilla/magpie-ultra-v0.1, only instructions rated as good or excellent were selected.
output_quality corresponds to the original dataset’s score_difference, which is the gap between instruct model and base model responses as evaluated by a reward model.
Please see the original dataset for details.
Source
Rows
argilla/magpie-ultra-v0.1
43923
Mxode/Magpie-Pro-10K-GPT4o-mini10000
Magpie-Llama-3.1-Pro-500K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-500K-Filtered.swallow-magpie-ultra-v0.1
📰 News
[07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs.
Dataset Summary
Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2.
The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.magpie
Magpie Collection
A collection of datasets where the prompts and the responses are generated by the models themselves from scratch (!)
Rows from the Magpie-Align/* datasets were included only if they were in English, had "good" question quality, and an answer quality score above five. Entries mentioning "Alibaba" in either the input or output were excluded.
Source
Rows
Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1
495048
HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered
328772… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/magpie.smoltalk-smol-magpie-ultra-no-refusals
SmolTalk Smol-Magpie-Ultra No Refusals
A Minos-cleaned version of HuggingFaceTB/smoltalk / smol-magpie-ultra for use as a neutral helpfulness SFT anchor.
Rows are removed when NousResearch/Minos-v1 classifies the conversation as a refusal. The original train/test split structure is preserved.
Cleaning version: minos-only-v1-2026-06-23
Counts
Split
Input rows
Kept rows
Dropped rows
train
409,537
408,447
1,090
test
21,555
21,488
67
Overall removal… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/smoltalk-smol-magpie-ultra-no-refusals.spoken-magpie-ja
Spoken-magpie
LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。
ある程度の話者多様性を持つように生成されています。
Respone Audioは500文字以下の場合にのみ生成されています。
NVIDIA H200を10枚を使用しvllmで推論しました。
Samples
最初の50サンプルを掲載します。
ID
Instruction
Instruction Audio
Response
Response Audio
0
カボチャを使ったスイーツのレシピをいくつか教えてください。
もちろんです、カボチャを使ったスイーツは秋にぴったりですね。以下にいくつかのレシピをご紹介します。1. カボチャのスフレパウンドケーキ- 材料:カボチャ 200g、生クリーム 50ml、牛乳 50ml、卵 3個、砂糖 100g、薄力粉 70g、バニラエッセンス 少々-… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-magpie-ja.Magpie-Air-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Air-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b.
This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Air-MT-300K-v0.1-ko.Magpie-Tanuki-8B-97k
Magpie-Tanuki-8B-97k
Magpieの手法をweblab-GENIAC/Tanuki-8B-dpo-v1.0に対して適用し作成した、97269件の日本語対話データセットです。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1
### Dataset Summary
`magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.Magpie-Llama-3.1-Pro-1M-v0.1-ko
일부 필드 번역이 안되어 재번역 예정입니다.
Translated Magpie-Align/Magpie-Llama-3.1-Pro-1M using nayohan/llama3-instrucTrans-enko-8b.
For this dataset, we only used data that is 5000 characters or less in length and has language of English.
Thanks for @Magpie-Align and @nayohan.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/Magpie-Llama-3.1-Pro-1M-v0.1-ko.sutra-magpie-sft
Sutra Magpie SFT Dataset
A high-quality dataset of 20,682 instruction-response pairs for supervised fine-tuning (SFT) of language models. Generated using seed prompts from the Sutra framework with Magpie-style response generation.
Dataset Description
This dataset provides diverse, high-quality instruction-response pairs suitable for training instruction-following language models.
Generation Method
Seed Prompts: Started with 30K diverse seed prompts from… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-magpie-sft.Magpie-Pro-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Pro-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b.
This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Pro-MT-300K-v0.1-ko.GLM-5.2-FP8-magpie-ultrachat
GLM-5.2-FP8 Regenerated Responses (Magpie + UltraChat mix)
A combined instruction-response dataset of 507,864 single-turn conversations. The
prompts are drawn from two public instruction datasets; the responses were freshly
regenerated with zai-org/GLM-5.2-FP8.
It was built as on-policy distillation data for training speculative-decoding drafts
(DFlash / DSpark) for GLM-5.2 — i.e. so the draft learns from GLM-5.2's own output
distribution — but it is a general-purpose GLM-5.2… See the full description on the dataset page: https://huggingface.co/datasets/mgoin/GLM-5.2-FP8-magpie-ultrachat.Magpie-Tanuki-8B-annotated-96k
Magpie-Tanuki-8B-annotated-96k
Magpieの手法をweblab-GENIAC/Tanuki-8B-dpo-v1.0に対して適用し作成したデータセットであるAratako/Magpie-Tanuki-8B-97kに対して、cyberagent/calm3-22b-chatを用いてinstructionに対して難易度、クオリティ、カテゴリをアノテーションしたデータセットです。
アノテーションのプロンプト
calm3によるアノテーションにはそれぞれ以下のプロンプトを利用しました。
難易度のアノテーション
# 指示
まず、与えられたユーザーの意図を特定し、その後、ユーザーのクエリの内容に基づいて難易度レベルをラベル付けしてください。
## ユーザーのクエリ
```
{input}
```
## 出力フォーマット
ユーザーのクエリに基づき、まずユーザーの意図を特定し、そのクエリを解決するために必要な知識を明示してください。
その後、難易度レベルを `very… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Magpie-Tanuki-8B-annotated-96k.magpie-sft-v1.0-dpo-judged
magpie-sft-v1.0-dpo-judged
概要
llm-jp/magpie-sft-v1.0を元に以下のような改変を加えて作成した日本語Preferenceデータセットです。
開発途中のモデルであるAratako/Llama-Gemma-2-27b-SFT-trial1を用いて回答を再生成
元データセットにあるQwen/Qwen2.5-32B-Instructの回答と再生成した回答の2つを並べ、google/gemma-2-27b-itによりどちらの回答の方が良いかをJudge
良いと判断された方の回答をchosenに、そうでない方の回答をrejectedに配置
ライセンス
本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。
META LLAMA 3.1 COMMUNITY LICENSEを継承します。
Gemma Terms of Useを継承します。
Qwen LICENSE… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/magpie-sft-v1.0-dpo-judged.Magpie-Llama-3.1-Pro-DPO-100K-v0.1-Magpie-Align
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/GenRM/Magpie-Llama-3.1-Pro-DPO-100K-v0.1-Magpie-Align.Synthetic-JP-Conversations-Magpie-Nemotron-4-10k
Synthetic-JP-Conversations-Magpie-Nemotron-4-10k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語instruction tuning用データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Magpie-Pro-10K-GPT4o-miniSynthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
swallow-gemma-magpie-v0.1
📰 News
[07/01/2025] Release of the first unfiltered version of the dataset containing 148k pairs.
Dataset Summary
Swallow-Gemma-Magpie-v0.1 is a synthetic instruction tuning dataset that consists of multiple category Japanese question-answering tasks.
It consists of 148k question-answering-samples, generated with google/gemma-2-27b-it.
Part of Swallow-Gemma-Magpie-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-gemma-magpie-v0.1.HQ-knowledgedistills-1.2M-magpieThis dataset is.an exact mix of 900k general qwen conversation with general questions, math, code and another 300k of Gemma 2 27B generations, for creative writing.
The dataset was made for "healing" pruned LLM's, especially ones based off of qwen2.5 series, as some conversations include the models saying who they are.
Unlike the previous 900K version, we also mixed in Gemma generations, to add more creative writing examples.
Many thanks to the magpie project for making this possible, this… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstackorg/HQ-knowledgedistills-1.2M-magpie.
