CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /magpie-ultra-v0.1 Dataset Card for magpie-ultra-v0.1 This dataset has been created with distilabel. 📰 News [26/11/2024] 🆕 New version of the dataset is out! magpie-ultra-v1.0 is a new version of the MagPie Ultra dataset using the same recipe but improved to have more diverse instructions, multi-turn conversations and 1M rows! [08/02/2024] Release of the first unfiltered version of the dataset containing 50K instruction-response pairs that can be used for SFT or… See the full description on the dataset page: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1.tabulartext-generation10K<n<100K221 likes2.6k downloads2y agoHugging Face02Magpie-Align /Magpie-Qwen2.5-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1.tabulartext-generation1M<n<10M19 likes2.2k downloads2y agoHugging Face03Magpie-Align /Magpie-Llama-3.1-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.4k downloads2y agoHugging Face04Magpie-Align /Magpie-Llama-3.1-Pro-MT-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.2k downloads2y agoHugging Face05Magpie-Align /Magpie-Llama-3.3-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-1M-v0.1.tabulartext-generation1M<n<10M5 likes370 downloads2y agoHugging Face06llm-jp /magpie-sft-v1.0 magpie-sft-v1.0 This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This is a dataset of instruction and response pairs created using the Magpie method. cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.texttext-generation100K<n<1M19 likes350 downloads2y agoHugging Face07Magpie-Align /Magpie-Llama-3.1-Pro-DPO-100K-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-DPO-100K-v0.1.texttext-generation100K<n<1M6 likes301 downloads2y agoHugging Face08Magpie-Align /Magpie-Llama-3.3-Pro-500K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-500K-Filtered.tabulartext-generation100K<n<1M3 likes299 downloads2y agoHugging Face09Magpie-Align /Magpie-Reasoning-V2-250K-CoT-Llama3 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Llama3.tabulartext-generation100K<n<1M11 likes189 downloads2y agoHugging Face10agentlans /small-magpie Smaller Magpie A collection of smaller Magpie datasets compared to agentlans/magpie. For argilla/magpie-ultra-v0.1, only instructions rated as good or excellent were selected. output_quality corresponds to the original dataset’s score_difference, which is the gap between instruct model and base model responses as evaluated by a reward model. Please see the original dataset for details. Source Rows argilla/magpie-ultra-v0.1 43923 Mxode/Magpie-Pro-10K-GPT4o-mini10000 texttext-generation100K<n<1M0 likes178 downloads10mo agoHugging Face11Magpie-Align /Magpie-Llama-3.1-Pro-500K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-500K-Filtered.tabulartext-generation100K<n<1M9 likes175 downloads2y agoHugging Face12tokyotech-llm /swallow-magpie-ultra-v0.1 📰 News [07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs. Dataset Summary Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2. The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.texttext-generation10K<n<100K5 likes151 downloads2y agoHugging Face13agentlans /magpie Magpie Collection A collection of datasets where the prompts and the responses are generated by the models themselves from scratch (!) Rows from the Magpie-Align/* datasets were included only if they were in English, had "good" question quality, and an answer quality score above five. Entries mentioning "Alibaba" in either the input or output were excluded. Source Rows Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1 495048 HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered 328772… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/magpie.texttext-generation1M<n<10M1 likes127 downloads8mo agoHugging Face14nchapman /smoltalk-smol-magpie-ultra-no-refusals SmolTalk Smol-Magpie-Ultra No Refusals A Minos-cleaned version of HuggingFaceTB/smoltalk / smol-magpie-ultra for use as a neutral helpfulness SFT anchor. Rows are removed when NousResearch/Minos-v1 classifies the conversation as a refusal. The original train/test split structure is preserved. Cleaning version: minos-only-v1-2026-06-23 Counts Split Input rows Kept rows Dropped rows train 409,537 408,447 1,090 test 21,555 21,488 67 Overall removal… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/smoltalk-smol-magpie-ultra-no-refusals.tabulartext-generation100K<n<1M1 likes114 downloads3mo agoHugging Face15Atotti /spoken-magpie-ja Spoken-magpie LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。 ある程度の話者多様性を持つように生成されています。 Respone Audioは500文字以下の場合にのみ生成されています。 NVIDIA H200を10枚を使用しvllmで推論しました。 Samples 最初の50サンプルを掲載します。 ID Instruction Instruction Audio Response Response Audio 0 カボチャを使ったスイーツのレシピをいくつか教えてください。 もちろんです、カボチャを使ったスイーツは秋にぴったりですね。以下にいくつかのレシピをご紹介します。1. カボチャのスフレパウンドケーキ- 材料:カボチャ 200g、生クリーム 50ml、牛乳 50ml、卵 3個、砂糖 100g、薄力粉 70g、バニラエッセンス 少々-… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-magpie-ja.audiotext-generation100K<n<1M1 likes112 downloads9mo agoHugging Face16nayohan /Magpie-Air-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Air-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b. This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered. @misc{xu2024magpie, title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing}, author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Air-MT-300K-v0.1-ko.texttext-generation100K<n<1M0 likes109 downloads2y agoHugging Face17Aratako /Magpie-Tanuki-8B-97k Magpie-Tanuki-8B-97k Magpieの手法をweblab-GENIAC/Tanuki-8B-dpo-v1.0に対して適用し作成した、97269件の日本語対話データセットです。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K12 likes101 downloads2y agoHugging Face181TuanPham /Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1 ### Dataset Summary `magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.textquestion-answering10K<n<100K1 likes100 downloads2y agoHugging Face19youjunhyeok /Magpie-Llama-3.1-Pro-1M-v0.1-ko 일부 필드 번역이 안되어 재번역 예정입니다. Translated Magpie-Align/Magpie-Llama-3.1-Pro-1M using nayohan/llama3-instrucTrans-enko-8b. For this dataset, we only used data that is 5000 characters or less in length and has language of English. Thanks for @Magpie-Align and @nayohan. @misc{xu2024magpie, title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing}, author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/Magpie-Llama-3.1-Pro-1M-v0.1-ko.tabulartext-generation100K<n<1M0 likes88 downloads2y agoHugging Face20codelion /sutra-magpie-sft Sutra Magpie SFT Dataset A high-quality dataset of 20,682 instruction-response pairs for supervised fine-tuning (SFT) of language models. Generated using seed prompts from the Sutra framework with Magpie-style response generation. Dataset Description This dataset provides diverse, high-quality instruction-response pairs suitable for training instruction-following language models. Generation Method Seed Prompts: Started with 30K diverse seed prompts from… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-magpie-sft.texttext-generation10K<n<100K2 likes88 downloads7mo agoHugging Face21nayohan /Magpie-Pro-MT-300K-v0.1-koTranslated Magpie-Align/Magpie-Pro-MT-300K-v0.1 using nayohan/llama3-instrucTrans-enko-8b. This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered. @misc{xu2024magpie, title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing}, author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Pro-MT-300K-v0.1-ko.texttext-generation100K<n<1M24 likes86 downloads2y agoHugging Face22mgoin /GLM-5.2-FP8-magpie-ultrachat GLM-5.2-FP8 Regenerated Responses (Magpie + UltraChat mix) A combined instruction-response dataset of 507,864 single-turn conversations. The prompts are drawn from two public instruction datasets; the responses were freshly regenerated with zai-org/GLM-5.2-FP8. It was built as on-policy distillation data for training speculative-decoding drafts (DFlash / DSpark) for GLM-5.2 — i.e. so the draft learns from GLM-5.2's own output distribution — but it is a general-purpose GLM-5.2… See the full description on the dataset page: https://huggingface.co/datasets/mgoin/GLM-5.2-FP8-magpie-ultrachat.texttext-generation100K<n<1M2 likes75 downloads3mo agoHugging Face23Aratako /Magpie-Tanuki-8B-annotated-96k Magpie-Tanuki-8B-annotated-96k Magpieの手法をweblab-GENIAC/Tanuki-8B-dpo-v1.0に対して適用し作成したデータセットであるAratako/Magpie-Tanuki-8B-97kに対して、cyberagent/calm3-22b-chatを用いてinstructionに対して難易度、クオリティ、カテゴリをアノテーションしたデータセットです。 アノテーションのプロンプト calm3によるアノテーションにはそれぞれ以下のプロンプトを利用しました。 難易度のアノテーション # 指示 まず、与えられたユーザーの意図を特定し、その後、ユーザーのクエリの内容に基づいて難易度レベルをラベル付けしてください。 ## ユーザーのクエリ ``` {input} ``` ## 出力フォーマット ユーザーのクエリに基づき、まずユーザーの意図を特定し、そのクエリを解決するために必要な知識を明示してください。 その後、難易度レベルを `very… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Magpie-Tanuki-8B-annotated-96k.texttext-generation10K<n<100K6 likes68 downloads2y agoHugging Face24Aratako /magpie-sft-v1.0-dpo-judged magpie-sft-v1.0-dpo-judged 概要 llm-jp/magpie-sft-v1.0を元に以下のような改変を加えて作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-SFT-trial1を用いて回答を再生成 元データセットにあるQwen/Qwen2.5-32B-Instructの回答と再生成した回答の2つを並べ、google/gemma-2-27b-itによりどちらの回答の方が良いかをJudge 良いと判断された方の回答をchosenに、そうでない方の回答をrejectedに配置 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1 COMMUNITY LICENSEを継承します。 Gemma Terms of Useを継承します。 Qwen LICENSE… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/magpie-sft-v1.0-dpo-judged.texttext-generation100K<n<1M0 likes65 downloads2y agoHugging Face25GenRM /Magpie-Llama-3.1-Pro-DPO-100K-v0.1-Magpie-Align Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/GenRM/Magpie-Llama-3.1-Pro-DPO-100K-v0.1-Magpie-Align.texttext-generation100K<n<1M0 likes63 downloads1y agoHugging Face26Aratako /Synthetic-JP-Conversations-Magpie-Nemotron-4-10k Synthetic-JP-Conversations-Magpie-Nemotron-4-10k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語instruction tuning用データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K11 likes58 downloads2y agoHugging Face27Mxode /Magpie-Pro-10K-GPT4o-minitexttext-generation10K<n<100K0 likes55 downloads1y agoHugging Face28Aratako /Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K6 likes54 downloads2y agoHugging Face29tokyotech-llm /swallow-gemma-magpie-v0.1 📰 News [07/01/2025] Release of the first unfiltered version of the dataset containing 148k pairs. Dataset Summary Swallow-Gemma-Magpie-v0.1 is a synthetic instruction tuning dataset that consists of multiple category Japanese question-answering tasks. It consists of 148k question-answering-samples, generated with google/gemma-2-27b-it. Part of Swallow-Gemma-Magpie-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-gemma-magpie-v0.1.texttext-generation100K<n<1M3 likes52 downloads2y agoHugging Face30Pinkstackorg /HQ-knowledgedistills-1.2M-magpieThis dataset is.an exact mix of 900k general qwen conversation with general questions, math, code and another 300k of Gemma 2 27B generations, for creative writing. The dataset was made for "healing" pruned LLM's, especially ones based off of qwen2.5 series, as some conversations include the models saying who they are. Unlike the previous 900K version, we also mixed in Gemma generations, to add more creative writing examples. Many thanks to the magpie project for making this possible, this… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstackorg/HQ-knowledgedistills-1.2M-magpie.texttext-generation1M<n<10M1 likes52 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.