CoolFace
Datasetpublic

nassimjp/Pashto-grammar-100

πŸ‡¦πŸ‡« Pashto Grammar 100 Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage. The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models. It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.

sourceHugging Facecc-by-4.0updated 9d agoView on Hugging Face
0likes37downloads
Dataset Card

πŸ‡¦πŸ‡« Pashto Grammar 100

Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.

The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.

It is intended as a small but high-quality resource for researchers and developers working on Pashto NLP and low-resource language modeling.


πŸ“Š Dataset Overview

PropertyValue
Dataset NamePashto Grammar 100
LanguagePashto (پښΨͺو)
Examples100
FormatJSONL
StructureInstruction / Question / Answer
DomainPashto Grammar
LicenseCC BY 4.0
TaskGrammar Analysis / Question Answering
Intended UseEducation, NLP, SFT, Evaluation

🎯 Purpose

The goal of Pashto Grammar 100 is to provide a compact collection of grammatical examples that can be used to teach or evaluate an AI model's ability to understand basic Pashto grammar.

The dataset focuses on areas such as:

  • β€”πŸ“ Sentence structure
  • β€”πŸ”€ Parts of speech
  • β€”πŸ§© Subject, object, and verb identification
  • —⏳ Verb tense
  • β€”πŸ”„ Verb usage
  • β€”πŸ‘₯ Nouns and pronouns
  • β€”πŸŽ― Adjectives
  • β€”πŸ› οΈ Adverbs
  • β€”πŸ”— Conjunctions
  • —❓ Interrogative sentences
  • β€”πŸš« Negative sentences
  • β€”πŸ“Œ Sentence patterns
  • β€”πŸ“š Basic grammatical rules
  • β€”πŸ‡¦πŸ‡« Standard Pashto usage

πŸ“‚ Dataset Format

The dataset uses a simple JSONL structure.

Each example contains a Pashto grammar question or sentence together with an appropriate explanation or answer.

Example:

json
{
  "messages": [
    {
      "role": "user",
      "content": "ΩΎΩ‡ دې Ψ¬Ω…Ω„Ω‡ کې فعل Ϊ©ΩˆΩ… دی؟ Β«Ψ§Ψ­Ω…Ψ― Ϊ©ΨͺΨ§Ψ¨ Ω„ΩˆΩ„ΩŠ.Β»"
    },
    {
      "role": "assistant",
      "content": "ΩΎΩ‡ دې Ψ¬Ω…Ω„Ω‡ کې Β«Ω„ΩˆΩ„ΩŠΒ» فعل دی."
    }
  ]
}

The messages structure makes the dataset suitable for modern conversational language models and common SFT training pipelines.


🧠 Grammar Coverage

The dataset provides examples covering several fundamental areas of Pashto grammar.

1. Sentence Structure

Examples help identify the basic structure of Pashto sentences, including:

  • β€”Subject
  • β€”Object
  • β€”Verb
  • β€”Sentence order
  • β€”Simple sentence construction

Pashto commonly follows an SOV (Subject–Object–Verb) pattern.

Example:

Ψ§Ψ­Ω…Ψ― Ϊ©ΨͺΨ§Ψ¨ Ω„ΩˆΩ„ΩŠ.
  • β€”Subject: Ψ§Ψ­Ω…Ψ―
  • β€”Object: Ϊ©ΨͺΨ§Ψ¨
  • β€”Verb: Ω„ΩˆΩ„ΩŠ

2. Verbs

The dataset includes examples involving Pashto verbs and their grammatical functions.

Examples include:

  • β€”Present tense
  • β€”Past tense
  • β€”Future tense
  • β€”Auxiliary usage
  • β€”Verb identification
  • β€”Verb agreement
  • β€”Simple verb constructions

3. Nouns

Examples demonstrate the identification and use of nouns in Pashto sentences.

Nouns may refer to:

  • β€”People
  • β€”Places
  • β€”Objects
  • β€”Animals
  • β€”Concepts

4. Pronouns

The dataset includes examples involving common Pashto pronouns and their grammatical roles.

Examples include:

  • β€”Ψ²Ω‡
  • β€”ΨͺΩ‡
  • β€”Ω‡ΨΊΩ‡
  • β€”Ω…ΩˆΪ–
  • β€”Ψͺاسو
  • β€”Ψ―ΩˆΫŒ

5. Adjectives

Examples demonstrate how adjectives describe nouns and how they function within Pashto sentences.

Example:

ΪšΪ©Ω„ΫŒ کور

Here:

  • β€”ΪšΪ©Ω„ΫŒ β†’ adjective
  • β€”Ϊ©ΩˆΨ± β†’ noun

6. Adverbs

Examples demonstrate words that describe how, when, or where an action takes place.


7. Negation

The dataset includes examples of negative sentence construction.

Example:

Ψ²Ω‡ ΪšΩˆΩˆΩ†ΪΩŠ ΨͺΩ‡ Ω†Ω‡ ځم.

The word Ω†Ω‡ is used to form the negative construction.


8. Questions

Examples include interrogative sentence structures.

Common question words include:

  • β€”Ϊ…ΩˆΪ©ΨŸ
  • β€”Ϊ…Ω‡ΨŸ
  • —چېرΨͺΩ‡ΨŸ
  • β€”Ϊ©Ω„Ω‡ΨŸ
  • β€”ΩˆΩ„ΫΨŸ
  • β€”Ϊ…Ω†Ϊ«Ω‡ΨŸ
  • β€”Ϊ©ΩˆΩ…ΨŸ

9. Tense

The dataset provides examples for recognizing different grammatical time references, including:

  • β€”Present
  • β€”Past
  • β€”Future

πŸ§ͺ Intended Uses

This dataset can be used for:

πŸ€– AI / NLP

  • β€”Pashto language model training
  • β€”Supervised fine-tuning (SFT)
  • β€”Instruction tuning
  • β€”Grammar evaluation
  • β€”Grammatical error analysis
  • β€”Pashto NLP research
  • β€”Low-resource language research

πŸŽ“ Education

  • β€”Pashto language learning
  • β€”Grammar tutoring
  • β€”Classroom exercises
  • β€”Self-study
  • β€”Automated grammar explanation

πŸ“Š Evaluation

The dataset can also be used as a small evaluation set for testing whether a language model can:

  • β€”Identify grammatical components
  • β€”Recognize verb tense
  • β€”Understand sentence structure
  • β€”Explain basic grammatical rules
  • β€”Produce grammatically appropriate Pashto

βš™οΈ Training Recommendations

Because this is a small 100-example dataset, it should not normally be used as a standalone training corpus for a large language model.

Instead, it can be combined with larger Pashto datasets for:

  • β€”SFT
  • β€”Instruction tuning
  • β€”Grammar specialization
  • β€”Domain adaptation
  • β€”Evaluation
  • β€”Final-stage refinement

The dataset is particularly useful as a small, focused grammar component inside a larger Pashto training mixture.


πŸ”¬ Research Applications

Possible research applications include:

  • β€”Pashto grammar modeling
  • β€”Low-resource language modeling
  • β€”Pashto instruction following
  • β€”Grammatical classification
  • β€”Grammar-aware LLM evaluation
  • β€”Pashto educational AI
  • β€”Multilingual model adaptation
  • β€”Pashto tokenizer and language-model research

🧹 Data Quality

The dataset is designed to remain:

  • β€”Focused on Pashto grammar
  • β€”Written in Pashto
  • β€”UTF-8 encoded
  • β€”Machine-learning friendly
  • β€”Simple to parse
  • β€”Suitable for conversational SFT formats

The examples are intended to emphasize grammatical understanding rather than general-domain knowledge.


⚠️ Limitations

This dataset contains only 100 examples and therefore should not be considered a comprehensive representation of the Pashto language.

It does not attempt to cover:

  • β€”All Pashto dialects
  • β€”Every grammatical construction
  • β€”Advanced morphology
  • β€”Complete verb paradigms
  • β€”Literary Pashto
  • β€”Specialized linguistic terminology
  • β€”All regional variations

For comprehensive Pashto grammar modeling, this dataset should be combined with larger and more diverse Pashto corpora.


πŸ“œ License

This dataset is released under the:

Creative Commons Attribution 4.0 International (CC BY 4.0)

You are free to:

  • β€”Share the dataset
  • β€”Adapt the dataset
  • β€”Use it for research
  • β€”Use it for machine-learning applications
  • β€”Use it in commercial projects

provided that appropriate attribution is given according to the CC BY 4.0 license.


πŸ‘€ Author

Nasibullah Nassim (Nassim JP)

Pashto AI Research & Development

πŸ‡―πŸ‡΅ Japan

Hugging Face:

nassimjp


🌐 Project

This dataset is part of the broader Pashto AI / iPashto.ai effort to develop open datasets and resources for Pashto language technology.

The broader goal is to improve the representation of Pashto language, grammar, reasoning, conversation, translation, and cultural knowledge in modern AI systems.


πŸ“– Citation

If you use this dataset in research or a project, please cite:

bibtex
@dataset{nassimjp_pashto_grammar_100_2026,
  author    = {Nasibullah Nassim},
  title     = {Pashto Grammar 100},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/nassimjp/Pashto-grammar-100},
  license   = {CC-BY-4.0}
}

🀝 Contributions

Contributions that improve Pashto language resources are welcome.

Possible future improvements include:

  • β€”Additional grammar examples
  • β€”More verb paradigms
  • β€”Morphological analysis
  • β€”Dialectal coverage
  • β€”Grammar correction examples
  • β€”Advanced syntax
  • β€”Linguistic annotations
  • β€”Native-speaker validation
  • β€”Larger evaluation sets

πŸ‡¦πŸ‡« Pashto AI

Pashto is a major language with millions of speakers, yet it remains significantly underrepresented in modern AI systems.

Small, focused datasets such as Pashto Grammar 100 can contribute to building better Pashto-capable language models by providing targeted linguistic knowledge that may be missing from large general-purpose corpora.

Ψ― پښΨͺو Ω„ΩΎΨ§Ψ±Ω‡ΨŒ Ω‡Ψ±Ω‡ ΩΎΨ§Ϊ©Ω‡ او Ϊ«ΩΌΩˆΨ±Ω‡ Ϊ©Ψ±ΪšΩ‡ ارزښΨͺ Ω„Ψ±ΩŠ. β€οΈπŸ‡¦πŸ‡«