CoolFace
Datasetpublic

nassimjp/pashto-dragon-1k-cot

Pashto-Dragon-1K-CoT Dataset Overview Pashto-Dragon-1K-CoT is a specialized reasoning dataset containing 1,000+ samples, meticulously translated into Pashto to facilitate the development of advanced Chain-of-Thought (CoT) capabilities in Pashto LLMs. This dataset is a high-quality derivative of the brendan-gho/qwen3b_paraphrased_dragon_cot. This repository is a core component of the iPashto.ai mission to move beyond simple web-scraping and focus on "Verified… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-dragon-1k-cot.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes8downloads
Dataset Card

Pashto-Dragon-1K-CoT Dataset

Overview

Pashto-Dragon-1K-CoT is a specialized reasoning dataset containing 1,000+ samples, meticulously translated into Pashto to facilitate the development of advanced Chain-of-Thought (CoT) capabilities in Pashto LLMs. This dataset is a high-quality derivative of the brendan-gho/qwen3b_paraphrased_dragon_cot.

This repository is a core component of the iPashto.ai mission to move beyond simple web-scraping and focus on "Verified Reasoning Data" for the Pashto language.

Dataset Structure

The dataset follows a structured CSV format with a focus on cross-lingual alignment:

  • prompt_en: The original complex reasoning query in English.
  • prompt_ps: The translated reasoning query in Pashto.
  • completion_en: The multi-step reasoning path and final answer in English.
  • completion_ps: The translated multi-step reasoning path and final answer in Pashto.

Strategic Importance

In the context of the iPashto.ai project, this dataset serves as a bridge between raw linguistic data and actual "Intelligence."

  • Reasoning-First: Unlike standard corpora, this focuses on how to solve a problem.
  • Tag Integrity: Special care has been taken to preserve <think> and <answer> tags, ensuring compatibility with modern reasoning architectures like DeepSeek-R1 style or OpenAI's O1-style training.
  • Consistency: Uses standardized Pashto terminology (e.g., preference for "ریږ" over "ریچھ") to ensure a native feel.

Methodology

  • Source: Qwen3B paraphrased "Dragon" CoT series.
  • Process: Automated neural translation followed by custom regex-based post-processing to fix tag formatting and linguistic inconsistencies.
  • Optimization: Designed for fine-tuning small to mid-sized models (e.g., Ghanam-1.B or Roshan).

About the Author & Project

Created and curated by Nassim الله (nassimjp), an IT Specialist and System Architect based in Saitama, Japan.

iPashto.ai is a long-term initiative dedicated to:

  1. 1.Building large-scale Pashto datasets (PashtoTM).
  2. 2.Developing native tokenizers and base models (Ghanam-1.B).
  3. 3.Implementing offline AI solutions for the Pashto language.

License

This dataset is provided under the Apache-2.0 license.