nassimjp/english_historical_quotes_in_pashto
📚 English Historical Quotes in Pashto — Chat Format Dataset A high‑quality, Pashto‑translated version of English Historical Quotes, converted into a chat‑style format suitable for training Pashto LLMs on quotation understanding, author attribution, and category‑based semantic reasoning. This dataset transforms each quote into: { "messages": [ {"role": "user", "content": "<Pashto Quote>"}, {"role": "assistant", "content": "لیکوال: <Author>\nکټګورۍ: <Categories>"} ] }… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_historical_quotes_in_pashto.
📚 English Historical Quotes in Pashto — Chat Format Dataset
A high‑quality, Pashto‑translated version of English Historical Quotes, converted into a chat‑style format suitable for training Pashto LLMs on quotation understanding, author attribution, and category‑based semantic reasoning.
This dataset transforms each quote into:
{
"messages": [
{"role": "user", "content": "<Pashto Quote>"},
{"role": "assistant", "content": "لیکوال: <Author>\nکټګورۍ: <Categories>"}
]
}🧩 Dataset Structure
Each row contains:
- id — unique string identifier
- messages — list of chat messages
- user → Pashto quote
- assistant → author + categories (Pashto)
Example:
{
"id": "001493",
"messages": [
{"role": "user", "content": "د خوښۍ حق بنسټیز دی."},
{"role": "assistant", "content": "لیکواله: انا پاولووا\nکټګورۍ: خوښۍ"}
]
}🎯 Purpose
This dataset is designed for:
- Pashto instruction‑tuning
- Pashto assistant alignment
- Quote‑based semantic reasoning
- Author attribution tasks
- Category‑based classification
- Lightweight conversational fine‑tuning
It is NOT intended for:
- Deep reasoning SFT
- Chain‑of‑thought training
- Multi‑turn conversation modeling
- Long‑context LLM training
🛠️ Source & Processing
Original dataset: m-ric/english_historical_quotes
Processing steps:
- Quotes translated into Pashto
- Author names normalized
- Categories translated and cleaned
- Converted into HF chat format
- UTF‑8 normalization
- Deterministic JSONL generation
📦 File Format
Dataset is provided as:
train.jsonl— main chat dataset- Each line is a complete chat sample
- Fully compatible with:
- HuggingFace
datasets - TRL
- OpenChat / ChatML loaders
- Talanda pipelines
🔍 Loading the Dataset
from datasets import load_dataset
ds = load_dataset("nassimjp/english_historical_quotes_in_pashto", split="train")
print(ds[0])📈 Recommended Use Cases
- Pashto assistant fine‑tuning
- Cultural & historical quote understanding
- Lightweight alignment tasks
- Category‑aware semantic modeling
- Author‑aware contextual reasoning
⚠️ Limitations
- Quotes are short → not suitable for long‑context training
- Assistant messages are metadata only → not reasoning
- No multi‑turn dialogue
- Not intended for CoT or deep instruction following
📜 License
This dataset follows the licensing terms of the original source dataset. Translations and chat formatting are provided by Spinzar Enterprises Inc. (Yūgen Kaisha).
🤝 Contributions
If you want to add:
- More quotes
- More categories
- Better translations
- Additional metadata
Feel free to open a PR or contact the maintainer.
🌟 Maintainer
Nassim Founder — Spinzar Enterprises Inc. (Yūgen Kaisha), Japan Pashto AI Research & Development Division
