CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llm-jp /magpie-sft-v1.0 magpie-sft-v1.0 This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This is a dataset of instruction and response pairs created using the Magpie method. cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.texttext-generation100K<n<1M19 likes347 downloads2y agoHugging Face02davanstrien /magpie-preference Dataset Card for Magpie Preference Dataset Dataset Description The Magpie Preference Dataset is a crowdsourced collection of human preferences on synthetic instruction-response pairs generated using the Magpie approach. This dataset is continuously updated through user interactions with the Magpie Preference Gradio Space. What is Magpie? Magpie is a very interesting new approach to creating synthetic data which doesn't require any seed data:… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/magpie-preference.textn<1K15 likes262 downloads2mo agoHugging Face03baseten-admin /magpie-qwen2.5-pro-1m-v0.1-Qwen3-235B-A22B-Instruct-2507-FP8-generatedtext1M<n<10M0 likes236 downloads11mo agoHugging Face04agentlans /small-magpie Smaller Magpie A collection of smaller Magpie datasets compared to agentlans/magpie. For argilla/magpie-ultra-v0.1, only instructions rated as good or excellent were selected. output_quality corresponds to the original dataset’s score_difference, which is the gap between instruct model and base model responses as evaluated by a reward model. Please see the original dataset for details. Source Rows argilla/magpie-ultra-v0.1 43923 Mxode/Magpie-Pro-10K-GPT4o-mini10000 texttext-generation100K<n<1M0 likes181 downloads10mo agoHugging Face05agentlans /argilla-magpie-ultratext1M<n<10M0 likes174 downloads8mo agoHugging Face06kgrabko /JiRack-Magpie-Pro-MT-300K_8k-Datasettext100K<n<1M0 likes169 downloads5mo agoHugging Face07yosefw /magpie-llama-3.2-1b-instructtext100K<n<1M0 likes159 downloads11d agoHugging Face08tokyotech-llm /swallow-magpie-ultra-v0.1 📰 News [07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs. Dataset Summary Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2. The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.texttext-generation10K<n<100K5 likes149 downloads2y agoHugging Face09baseten-admin /gpt-oss120b-generated-magpie-1m-v0.1text100K<n<1M2 likes145 downloads1y agoHugging Face10agentlans /magpie Magpie Collection A collection of datasets where the prompts and the responses are generated by the models themselves from scratch (!) Rows from the Magpie-Align/* datasets were included only if they were in English, had "good" question quality, and an answer quality score above five. Entries mentioning "Alibaba" in either the input or output were excluded. Source Rows Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1 495048 HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered 328772… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/magpie.texttext-generation1M<n<10M1 likes121 downloads8mo agoHugging Face11SillyTilly /Magpie-Pro-DPO-200K-JSONLtabular100K<n<1M5 likes114 downloads2y agoHugging Face12nyu-dice-lab /lm-eval-results-Magpie-Align-Llama-3-8B-OpenHermes-243K-private Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-OpenHermes-243K Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-OpenHermes-243K The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-OpenHermes-243K-private.tabular100K<n<1M0 likes103 downloads2y agoHugging Face131TuanPham /Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1 ### Dataset Summary `magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.textquestion-answering10K<n<100K1 likes99 downloads2y agoHugging Face14bofenghuang /magpie-frtabular1M<n<10M0 likes97 downloads2y agoHugging Face15usable-japanese-llm /sft-magpietext10K<n<100K1 likes91 downloads2y agoHugging Face16mgoin /GLM-5.2-FP8-magpie-ultrachat GLM-5.2-FP8 Regenerated Responses (Magpie + UltraChat mix) A combined instruction-response dataset of 507,864 single-turn conversations. The prompts are drawn from two public instruction datasets; the responses were freshly regenerated with zai-org/GLM-5.2-FP8. It was built as on-policy distillation data for training speculative-decoding drafts (DFlash / DSpark) for GLM-5.2 — i.e. so the draft learns from GLM-5.2's own output distribution — but it is a general-purpose GLM-5.2… See the full description on the dataset page: https://huggingface.co/datasets/mgoin/GLM-5.2-FP8-magpie-ultrachat.texttext-generation100K<n<1M2 likes78 downloads3mo agoHugging Face17nyu-dice-lab /lm-eval-results-Magpie-Align-Llama-3-8B-Tulu-330K-private Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-Tulu-330K Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-Tulu-330K The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-Tulu-330K-private.tabular100K<n<1M0 likes61 downloads2y agoHugging Face18Aratako /Magpie-Tanuki-Qwen2.5-72B-Answered Magpie-Tanuki-Qwen2.5-72B-Answered Aratako/Magpie-Tanuki-8B-annotated-96kからinput_qualityがexcellentのものを抽出し、それに対してQwen/Qwen2.5-72B-Instructで回答の再生成を行ったデータセットです。 ライセンス 基本的にはApache 2.0に準じますが、Qwen Licenseの影響を受けるため、このデータセットを使ってモデルを学習する際はこのライセンスの制約に従ってください。 text10K<n<100K1 likes59 downloads2y agoHugging Face19nyu-dice-lab /lm-eval-results-Magpie-Align-Llama-3-8B-WizardLM-196K-private Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-WizardLM-196K Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-WizardLM-196K The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-WizardLM-196K-private.tabular100K<n<1M0 likes59 downloads2y agoHugging Face20Aratako /Synthetic-JP-Conversations-Magpie-Nemotron-4-10k Synthetic-JP-Conversations-Magpie-Nemotron-4-10k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語instruction tuning用データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K11 likes56 downloads2y agoHugging Face21tokyotech-llm /swallow-gemma-magpie-v0.1 📰 News [07/01/2025] Release of the first unfiltered version of the dataset containing 148k pairs. Dataset Summary Swallow-Gemma-Magpie-v0.1 is a synthetic instruction tuning dataset that consists of multiple category Japanese question-answering tasks. It consists of 148k question-answering-samples, generated with google/gemma-2-27b-it. Part of Swallow-Gemma-Magpie-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-gemma-magpie-v0.1.texttext-generation100K<n<1M3 likes56 downloads2y agoHugging Face22Aratako /Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K6 likes55 downloads2y agoHugging Face23Mxode /Magpie-Pro-10K-GPT4o-minitexttext-generation10K<n<100K0 likes55 downloads1y agoHugging Face24nyu-dice-lab /lm-eval-results-Magpie-Align-Llama-3-8B-WildChat-private Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-WildChat Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-WildChat The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-WildChat-private.tabular100K<n<1M0 likes55 downloads2y agoHugging Face25Aratako /Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語のコーディング用対話データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K1 likes47 downloads2y agoHugging Face26DeL-TaiseiOzaki /magpie-qwen2.5-32B-10K-ja 合成日本語指示データセット 概要 このデータセットは、大規模言語モデル(Qwen2.5-32B-instruct)を用いて自動生成された日本語の指示とそれに対する応答のコレクションです。データセットは指示付与型のタスクのための学習や評価に使用することを目的としています。 データセット仕様 サンプル数: 20,000 言語: 日本語 フォーマット: JSON ライセンス: Apache-2.0 サイズカテゴリ: 10K<n<100K 生成方法 データセットは以下のプロセスを通じて生成されました: Qwen2.5-32B Instructモデルを使用 各サンプルは3段階のプロセスで生成: a) 指示文の生成 b) Chain-of-Thought (CoT) 応答の生成 (一部のデータには含まれない) c) 最終的な応答のself-refine 生成の多様性を向上させるため、10種類のペルソナからランダムに1つ選んでシステムプロンプトに入力 詳細な生成プロセスはこちらをご覧ください。… See the full description on the dataset page: https://huggingface.co/datasets/DeL-TaiseiOzaki/magpie-qwen2.5-32B-10K-ja.text10K<n<100K0 likes46 downloads2y agoHugging Face27Aratako /Synthetic-JP-EN-Coding-Dataset-Magpie-69k Synthetic-JP-EN-Coding-Dataset-Magpie-69k Magpieの手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。 作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1 cyberagent/calm3-22b-chat データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K9 likes44 downloads2y agoHugging Face28open-llm-leaderboard /Magpie-Align__Llama-3.1-8B-Magpie-Align-SFT-v0.1-detailsgated Dataset Card for Evaluation run of Magpie-Align/Llama-3.1-8B-Magpie-Align-SFT-v0.1 Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3.1-8B-Magpie-Align-SFT-v0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Magpie-Align__Llama-3.1-8B-Magpie-Align-SFT-v0.1-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face29HiTZ /Magpie-Llama-3.1-8B-Instruct-Filtered-1MDataset generated using meta-llama/Llama-3.1-8B-Instruct with the MAGPIE codebase. The unfiltered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered Filter criteria min_repetition = 100 def test_no_repetition(text: str): # Count the frequency of each word in the text word_count = Counter(text.split()) # Check if any word appears more than min_repetition times return all(count <= min_repetition for count in word_count.values()) def… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered-1M.text100K<n<1M0 likes34 downloads1y agoHugging Face30OpenLLM-Ro /ro_sft_magpie_mt Dataset Description Magpie is a data synthesis pipeline that generate high-quality alignment data. Magpie-Pro-MT dataset contains 300k instruction-following data generated with Llama3.1-70B. Here we provide the Romanian translation of the Magpie-Pro-MT dataset, translated with GPT-4o mini. This dataset represents a next step of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_magpie_mt.text100K<n<1M0 likes33 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.