datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
magpie-sft-v1.0
magpie-sft-v1.0
This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This is a dataset of instruction and response pairs created using the Magpie method.
cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.magpie-preference
Dataset Card for Magpie Preference Dataset
Dataset Description
The Magpie Preference Dataset is a crowdsourced collection of human preferences on synthetic instruction-response pairs generated using the Magpie approach.
This dataset is continuously updated through user interactions with the Magpie Preference Gradio Space.
What is Magpie?
Magpie is a very interesting new approach to creating synthetic data which doesn't require any seed data:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/magpie-preference.magpie-qwen2.5-pro-1m-v0.1-Qwen3-235B-A22B-Instruct-2507-FP8-generatedsmall-magpie
Smaller Magpie
A collection of smaller Magpie datasets compared to agentlans/magpie.
For argilla/magpie-ultra-v0.1, only instructions rated as good or excellent were selected.
output_quality corresponds to the original dataset’s score_difference, which is the gap between instruct model and base model responses as evaluated by a reward model.
Please see the original dataset for details.
Source
Rows
argilla/magpie-ultra-v0.1
43923
Mxode/Magpie-Pro-10K-GPT4o-mini10000
argilla-magpie-ultraJiRack-Magpie-Pro-MT-300K_8k-Datasetmagpie-llama-3.2-1b-instructswallow-magpie-ultra-v0.1
📰 News
[07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs.
Dataset Summary
Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2.
The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.gpt-oss120b-generated-magpie-1m-v0.1magpie
Magpie Collection
A collection of datasets where the prompts and the responses are generated by the models themselves from scratch (!)
Rows from the Magpie-Align/* datasets were included only if they were in English, had "good" question quality, and an answer quality score above five. Entries mentioning "Alibaba" in either the input or output were excluded.
Source
Rows
Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1
495048
HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered
328772… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/magpie.Magpie-Pro-DPO-200K-JSONLlm-eval-results-Magpie-Align-Llama-3-8B-OpenHermes-243K-private
Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-OpenHermes-243K
Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-OpenHermes-243K
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-OpenHermes-243K-private.Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1
### Dataset Summary
`magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.magpie-frsft-magpieGLM-5.2-FP8-magpie-ultrachat
GLM-5.2-FP8 Regenerated Responses (Magpie + UltraChat mix)
A combined instruction-response dataset of 507,864 single-turn conversations. The
prompts are drawn from two public instruction datasets; the responses were freshly
regenerated with zai-org/GLM-5.2-FP8.
It was built as on-policy distillation data for training speculative-decoding drafts
(DFlash / DSpark) for GLM-5.2 — i.e. so the draft learns from GLM-5.2's own output
distribution — but it is a general-purpose GLM-5.2… See the full description on the dataset page: https://huggingface.co/datasets/mgoin/GLM-5.2-FP8-magpie-ultrachat.lm-eval-results-Magpie-Align-Llama-3-8B-Tulu-330K-private
Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-Tulu-330K
Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-Tulu-330K
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-Tulu-330K-private.Magpie-Tanuki-Qwen2.5-72B-Answered
Magpie-Tanuki-Qwen2.5-72B-Answered
Aratako/Magpie-Tanuki-8B-annotated-96kからinput_qualityがexcellentのものを抽出し、それに対してQwen/Qwen2.5-72B-Instructで回答の再生成を行ったデータセットです。
ライセンス
基本的にはApache 2.0に準じますが、Qwen Licenseの影響を受けるため、このデータセットを使ってモデルを学習する際はこのライセンスの制約に従ってください。
lm-eval-results-Magpie-Align-Llama-3-8B-WizardLM-196K-private
Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-WizardLM-196K
Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-WizardLM-196K
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-WizardLM-196K-private.Synthetic-JP-Conversations-Magpie-Nemotron-4-10k
Synthetic-JP-Conversations-Magpie-Nemotron-4-10k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語instruction tuning用データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
swallow-gemma-magpie-v0.1
📰 News
[07/01/2025] Release of the first unfiltered version of the dataset containing 148k pairs.
Dataset Summary
Swallow-Gemma-Magpie-v0.1 is a synthetic instruction tuning dataset that consists of multiple category Japanese question-answering tasks.
It consists of 148k question-answering-samples, generated with google/gemma-2-27b-it.
Part of Swallow-Gemma-Magpie-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-gemma-magpie-v0.1.Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Magpie-Pro-10K-GPT4o-minilm-eval-results-Magpie-Align-Llama-3-8B-WildChat-private
Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-WildChat
Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-WildChat
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-WildChat-private.Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語のコーディング用対話データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
magpie-qwen2.5-32B-10K-ja
合成日本語指示データセット
概要
このデータセットは、大規模言語モデル(Qwen2.5-32B-instruct)を用いて自動生成された日本語の指示とそれに対する応答のコレクションです。データセットは指示付与型のタスクのための学習や評価に使用することを目的としています。
データセット仕様
サンプル数: 20,000
言語: 日本語
フォーマット: JSON
ライセンス: Apache-2.0
サイズカテゴリ: 10K<n<100K
生成方法
データセットは以下のプロセスを通じて生成されました:
Qwen2.5-32B Instructモデルを使用
各サンプルは3段階のプロセスで生成:
a) 指示文の生成
b) Chain-of-Thought (CoT) 応答の生成 (一部のデータには含まれない)
c) 最終的な応答のself-refine
生成の多様性を向上させるため、10種類のペルソナからランダムに1つ選んでシステムプロンプトに入力
詳細な生成プロセスはこちらをご覧ください。… See the full description on the dataset page: https://huggingface.co/datasets/DeL-TaiseiOzaki/magpie-qwen2.5-32B-10K-ja.Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Magpieの手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。
作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1
cyberagent/calm3-22b-chat
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Magpie-Align__Llama-3.1-8B-Magpie-Align-SFT-v0.1-details
Dataset Card for Evaluation run of Magpie-Align/Llama-3.1-8B-Magpie-Align-SFT-v0.1
Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3.1-8B-Magpie-Align-SFT-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Magpie-Align__Llama-3.1-8B-Magpie-Align-SFT-v0.1-details.Magpie-Llama-3.1-8B-Instruct-Filtered-1MDataset generated using meta-llama/Llama-3.1-8B-Instruct with the MAGPIE codebase.
The unfiltered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered
Filter criteria
min_repetition = 100
def test_no_repetition(text: str):
# Count the frequency of each word in the text
word_count = Counter(text.split())
# Check if any word appears more than min_repetition times
return all(count <= min_repetition for count in word_count.values())
def… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered-1M.ro_sft_magpie_mt
Dataset Description
Magpie is a data synthesis pipeline that generate high-quality alignment data. Magpie-Pro-MT dataset contains 300k instruction-following data generated with Llama3.1-70B.
Here we provide the Romanian translation of the Magpie-Pro-MT dataset, translated with GPT-4o mini.
This dataset represents a next step of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_magpie_mt.
