CoolFace
Datasetpublic

HiTZ/magpie-en-eu-reasoning-instructions-qwen3

Dataset Card for magpie-en-eu-reasoning-instructions-qwen3 Dataset Summary The magpie-en-eu-reasoning-instructions-qwen3 dataset is a large-scale, high-quality, bilingual instruction and preference dataset developed by the HiTZ Center. It is specifically tailored for training, aligning, and evaluating reasoning-focused Large Language Models (LLMs) in both English and Basque (Euskera). Built using the self-synthesizing Magpie methodology, the dataset contains a… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/magpie-en-eu-reasoning-instructions-qwen3.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
1likes377downloads
Dataset Card

Dataset Card for magpie-en-eu-reasoning-instructions-qwen3

Dataset Summary

The magpie-en-eu-reasoning-instructions-qwen3 dataset is a large-scale, high-quality, bilingual instruction and preference dataset developed by the HiTZ Center. It is specifically tailored for training, aligning, and evaluating reasoning-focused Large Language Models (LLMs) in both English and Basque (Euskera).

Built using the self-synthesizing Magpie methodology, the dataset contains a total of 1,000,000 instructions (500,000 in English and 500,000 in Basque). For preference alignment, every instruction is paired with two responses (one chosen, one rejected), totaling 2,000,000 responses. Uniquely, as this dataset targets reasoning models, both responses include the explicit step-by-step reasoning traces used to arrive at the final answer.

Dataset Details

  • —Developed by: HiTZ Center (Basque Center for Language Technology)
  • —Language(s) (NLP): English (en), Basque (eu)
  • —License: Permissive Open-Source Licenses
  • —Repository: HuggingFace Dataset
  • —Total Instructions: 1,000,000
  • —Total Responses/Preferences: 2,000,000 paired combinations
  • —Primary Models Used: Qwen3-32B, Qwen3-235B-A22B, and a Basque-adapted Qwen-32B

Dataset Structure

Data Instances

Each instance in the dataset represents a complete preference pair containing an instruction, a domain category, the source language, and two distinct reasoning-response pairs (chosen vs. rejected).

Data Fields

  • —instruction (string): The prompt or user instruction generated via Magpie.
  • —category (string): The thematic domain of the instruction (one of 14 categories).
  • —language (string): The language indicator (en for English, eu for Basque).
  • —chosen_response (string): The higher-quality answer generated by the larger model (Qwen3-235B-A22B).
  • —chosen_reasoning (string): The detailed internal reasoning trace/thought process behind the chosen response.
  • —rejected_response (string): The lower-quality answer generated by the smaller model (Qwen3-32B).
  • —rejected_reasoning (string): The internal reasoning trace/thought process behind the rejected response.

Category Distribution

Instructions are systematically curated across 14 distinct categories using specialized system prompts:

  1. 1.Translation
  2. 2.Science
  3. 3.Safety
  4. 4.Friendly Chat
  5. 5.Creative Writing
  6. 6.Business
  7. 7.Generic
  8. 8.Debugging
  9. 9.Mathematics
  10. 10.Reasoning
  11. 11.Education
  12. 12.Programming
  13. 13.Data Analysis
  14. 14.Documentation

Dataset Creation

1. Instruction Generation & Quality Control

The instructions were generated autonomously following the Magpie methodology (Xu et al., 2025). Instead of using explicit human seed prompts, instructions are extracted directly by sampling from the Qwen3-32B model starting from specific system prompts. This allows the model to map out instructions natively according to its pre-trained data distribution.

To guarantee high standard outputs, all instructions underwent a rigorous automated multi-dimensional evaluation. The generating model evaluated each prompt based on:

  • —Clarity
  • —Completeness
  • —Specificity
  • —Complexity
  • —Actionability
  • —Coherence

A global composite score was computed from these traits, and any instruction failing to meet the quality threshold was discarded.

2. Response & Preference Generation

For every filtered instruction, answers were generated simultaneously by two models from the same architectural family but possessing different capacities:

  • —Chosen (Preferred): Generated by Qwen3-235B-A22B
  • —Rejected (Less Preferred): Generated by Qwen3-32B

Because both models are reasoning-centric architectures, their internal generation traces (thinking steps) were recorded and stored in the dataset (chosen_reasoning and rejected_reasoning) alongside the final textual responses. This setup provides an ideal playground for Direct Preference Optimization (DPO), Odds Ratio Preference Optimization (ORPO), or RLHF tailored for reasoning capabilities.

3. Basque Localization & Translation

To build the corresponding Basque parallel corpus, the high-quality English instructions, reasonings, and responses were processed using a specialized Basque-adapted Qwen-32B model. This specific model was fine-tuned to properly grasp the unique morphological, syntactic, and cultural intricacies of the Basque language.

A subsequent manual evaluation on a randomized representative sample was conducted by human experts to ensure the structural integrity, natural flow, and semantic accuracy of the Basque translations matched the English baseline.


Use Cases & Intentions

This dataset is primarily built for:

  • —Preference Alignment: Training reward models or direct optimization algorithms (DPO, IPO, KTO, ORPO).
  • —Reasoning Capabilities: Teaching models how to think before answering by utilizing the embedded reasoning traces.
  • —Bilingual Capabilities: Enhancing the instruction-following and complex problem-solving capabilities of Basque and English language models.

Acknowledgements

This dataset was produced by the HiTZ Center as part of its ongoing research into low-resource language enhancement, language model alignment methodologies, and robust evaluation paradigms.

Funding

This work is funded by the Basque Government (IKER-GAITU project) and the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project ILENIA with reference 2022/TL22/00215335 and within the framework of the project Desarrollo de Modelos ALIA.