nassimjp/goal-pashto-chat-sharegpt-5GB
📄 goal-pashto-chat-sharegpt-5GB — Pashto ShareGPT‑Style Chat Dataset A large‑scale, high‑quality Pashto conversational dataset designed for instruction‑tuning, dialogue modeling, and LLM alignment.This dataset contains ~5GB of multi‑turn Pashto conversations inspired by ShareGPT, covering reasoning, advice, education, culture, and general knowledge. 📌 Dataset Summary goal-pashto-chat-sharegpt-5GB is a curated collection of Pashto user–assistant conversations… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB.
📄 goal-pashto-chat-sharegpt-5GB — Pashto ShareGPT‑Style Chat Dataset
A large‑scale, high‑quality Pashto conversational dataset designed for instruction‑tuning, dialogue modeling, and LLM alignment. This dataset contains ~5GB of multi‑turn Pashto conversations inspired by ShareGPT, covering reasoning, advice, education, culture, and general knowledge.
📌 Dataset Summary
goal-pashto-chat-sharegpt-5GB is a curated collection of Pashto user–assistant conversations following the standard chat schema:
{
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]
}
The dataset is optimized for:
- Pashto chatbots
- Instruction‑following LLMs
- Multi‑turn dialogue agents
- Cultural & linguistic alignment
- Long‑context reasoning models (4k–16k)
📂 Data Structure
Each entry contains:
Example
{
"messages": [
{
"role": "user",
"content": "په افغانستان کې د ژمي موسم څنګه وي؟"
},
{
"role": "assistant",
"content": "په ډېرو سیمو کې ژمی یخ او واورین وي، په ځانګړي ډول په غرنیو ولایتونو کې."
}
]
}
📊 Token Length Distribution
This dataset contains short, medium, long, and very long conversations:
This makes the dataset ideal for training 8k–16k context Pashto LLMs.
🧠 Use Cases
- Pashto conversational AI
- Chatbot fine‑tuning (SFT)
- Cultural alignment for LLMs
- Long‑context reasoning
- Multi‑turn dialogue generation
- Low‑resource NLP research
🛠 Training Format
Compatible with:
- LLaMA‑3 / LLaMA‑2 chat templates
- Qwen chat template
- Unsloth SFT pipelines
- Talanda data‑engineering framework
Example training snippet
tokenizer.apply_chat_template(
example["messages"],
tokenize=True,
add_generation_prompt=False
)
📥 Loading the Dataset
from datasets import load_dataset
ds = load_dataset("nassimjp/goal-pashto-chat-sharegpt-5GB")
📚 Citation
@dataset{nassimjp2026goalpashtochat,
title = {Goal Pashto Chat ShareGPT 5GB},
author = {Nassimjp},
year = {2026},
publisher = {Hugging Face Datasets},
url = {[https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB](https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB)},
note = {A large-scale Pashto conversational dataset for instruction tuning and chat-based LLM training.}
}
📜 License
This dataset is released under a non‑commercial research license.
Please review the license terms before using it in production systems.
🙌 Acknowledgements
This dataset is part of the broader effort to build high‑quality Pashto NLP resources and support low‑resource language AI development.
