CoolFace
Datasetpublic

nassimjp/english_historical_quotes_in_pashto

📚 English Historical Quotes in Pashto — Chat Format Dataset A high‑quality, Pashto‑translated version of English Historical Quotes, converted into a chat‑style format suitable for training Pashto LLMs on quotation understanding, author attribution, and category‑based semantic reasoning. This dataset transforms each quote into: { "messages": [ {"role": "user", "content": "<Pashto Quote>"}, {"role": "assistant", "content": "لیکوال: <Author>\nکټګورۍ: <Categories>"} ] }… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_historical_quotes_in_pashto.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes13downloads
Dataset Card

📚 English Historical Quotes in Pashto — Chat Format Dataset

A high‑quality, Pashto‑translated version of English Historical Quotes, converted into a chat‑style format suitable for training Pashto LLMs on quotation understanding, author attribution, and category‑based semantic reasoning.

This dataset transforms each quote into:

json
{
  "messages": [
    {"role": "user", "content": "<Pashto Quote>"},
    {"role": "assistant", "content": "لیکوال: <Author>\nکټګورۍ: <Categories>"}
  ]
}

🧩 Dataset Structure

Each row contains:

  • —id — unique string identifier
  • —messages — list of chat messages
  • —user → Pashto quote
  • —assistant → author + categories (Pashto)

Example:

json
{
  "id": "001493",
  "messages": [
    {"role": "user", "content": "د خوښۍ حق بنسټیز دی."},
    {"role": "assistant", "content": "لیکواله: انا پاولووا\nکټګورۍ: خوښۍ"}
  ]
}

🎯 Purpose

This dataset is designed for:

  • —Pashto instruction‑tuning
  • —Pashto assistant alignment
  • —Quote‑based semantic reasoning
  • —Author attribution tasks
  • —Category‑based classification
  • —Lightweight conversational fine‑tuning

It is NOT intended for:

  • —Deep reasoning SFT
  • —Chain‑of‑thought training
  • —Multi‑turn conversation modeling
  • —Long‑context LLM training

🛠️ Source & Processing

Original dataset: m-ric/english_historical_quotes

Processing steps:

  1. 1.Quotes translated into Pashto
  2. 2.Author names normalized
  3. 3.Categories translated and cleaned
  4. 4.Converted into HF chat format
  5. 5.UTF‑8 normalization
  6. 6.Deterministic JSONL generation

📦 File Format

Dataset is provided as:

  • —train.jsonl — main chat dataset
  • —Each line is a complete chat sample
  • —Fully compatible with:
  • —HuggingFace datasets
  • —TRL
  • —OpenChat / ChatML loaders
  • —Talanda pipelines

🔍 Loading the Dataset

python
from datasets import load_dataset

ds = load_dataset("nassimjp/english_historical_quotes_in_pashto", split="train")

print(ds[0])

📈 Recommended Use Cases

  • —Pashto assistant fine‑tuning
  • —Cultural & historical quote understanding
  • —Lightweight alignment tasks
  • —Category‑aware semantic modeling
  • —Author‑aware contextual reasoning

⚠️ Limitations

  • —Quotes are short → not suitable for long‑context training
  • —Assistant messages are metadata only → not reasoning
  • —No multi‑turn dialogue
  • —Not intended for CoT or deep instruction following

📜 License

This dataset follows the licensing terms of the original source dataset. Translations and chat formatting are provided by Spinzar Enterprises Inc. (Yūgen Kaisha).


🤝 Contributions

If you want to add:

  • —More quotes
  • —More categories
  • —Better translations
  • —Additional metadata

Feel free to open a PR or contact the maintainer.


🌟 Maintainer

Nassim Founder — Spinzar Enterprises Inc. (Yūgen Kaisha), Japan Pashto AI Research & Development Division