CoolFace
Datasetpublic

nassimjp/goal-pashto-chat-sharegpt-5GB

📄 goal-pashto-chat-sharegpt-5GB — Pashto ShareGPT‑Style Chat Dataset A large‑scale, high‑quality Pashto conversational dataset designed for instruction‑tuning, dialogue modeling, and LLM alignment.This dataset contains ~5GB of multi‑turn Pashto conversations inspired by ShareGPT, covering reasoning, advice, education, culture, and general knowledge. 📌 Dataset Summary goal-pashto-chat-sharegpt-5GB is a curated collection of Pashto user–assistant conversations… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes13downloads
Dataset Card

📄 goal-pashto-chat-sharegpt-5GB — Pashto ShareGPT‑Style Chat Dataset

A large‑scale, high‑quality Pashto conversational dataset designed for instruction‑tuning, dialogue modeling, and LLM alignment. This dataset contains ~5GB of multi‑turn Pashto conversations inspired by ShareGPT, covering reasoning, advice, education, culture, and general knowledge.


📌 Dataset Summary

goal-pashto-chat-sharegpt-5GB is a curated collection of Pashto user–assistant conversations following the standard chat schema:

json
{
  "messages": [
    {"role": "user", "content": "..."},
    {"role": "assistant", "content": "..."}
  ]
}

The dataset is optimized for:

  • —Pashto chatbots
  • —Instruction‑following LLMs
  • —Multi‑turn dialogue agents
  • —Cultural & linguistic alignment
  • —Long‑context reasoning models (4k–16k)

📂 Data Structure

Each entry contains:

FieldDescription
messagesList of chat turns
role"user" or "assistant"
contentPashto text of the message

Example

json
{
  "messages": [
    {
      "role": "user",
      "content": "په افغانستان کې د ژمي موسم څنګه وي؟"
    },
    {
      "role": "assistant",
      "content": "په ډېرو سیمو کې ژمی یخ او واورین وي، په ځانګړي ډول په غرنیو ولایتونو کې."
    }
  ]
}

📊 Token Length Distribution

This dataset contains short, medium, long, and very long conversations:

BucketCount
≤51245
≤1024341
≤2048425
≤4096340
≤8192included
≤16384included
>16384included

This makes the dataset ideal for training 8k–16k context Pashto LLMs.


🧠 Use Cases

  • —Pashto conversational AI
  • —Chatbot fine‑tuning (SFT)
  • —Cultural alignment for LLMs
  • —Long‑context reasoning
  • —Multi‑turn dialogue generation
  • —Low‑resource NLP research

🛠 Training Format

Compatible with:

  • —LLaMA‑3 / LLaMA‑2 chat templates
  • —Qwen chat template
  • —Unsloth SFT pipelines
  • —Talanda data‑engineering framework

Example training snippet

python
tokenizer.apply_chat_template(
    example["messages"],
    tokenize=True,
    add_generation_prompt=False
)

📥 Loading the Dataset

python
from datasets import load_dataset
ds = load_dataset("nassimjp/goal-pashto-chat-sharegpt-5GB")

📚 Citation

bibtex
@dataset{nassimjp2026goalpashtochat,
  title        = {Goal Pashto Chat ShareGPT 5GB},
  author       = {Nassimjp},
  year         = {2026},
  publisher    = {Hugging Face Datasets},
  url          = {[https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB](https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB)},
  note         = {A large-scale Pashto conversational dataset for instruction tuning and chat-based LLM training.}
}

📜 License

This dataset is released under a non‑commercial research license.

Please review the license terms before using it in production systems.


🙌 Acknowledgements

This dataset is part of the broader effort to build high‑quality Pashto NLP resources and support low‑resource language AI development.