nassimjp/Pashto-grammar-100
π¦π« Pashto Grammar 100 Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage. The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models. It is intended as a small but high-quality resource for researchers and developersβ¦ See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.
π¦π« Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers working on Pashto NLP and low-resource language modeling.
π Dataset Overview
π― Purpose
The goal of Pashto Grammar 100 is to provide a compact collection of grammatical examples that can be used to teach or evaluate an AI model's ability to understand basic Pashto grammar.
The dataset focuses on areas such as:
- π Sentence structure
- π€ Parts of speech
- π§© Subject, object, and verb identification
- β³ Verb tense
- π Verb usage
- π₯ Nouns and pronouns
- π― Adjectives
- π οΈ Adverbs
- π Conjunctions
- β Interrogative sentences
- π« Negative sentences
- π Sentence patterns
- π Basic grammatical rules
- π¦π« Standard Pashto usage
π Dataset Format
The dataset uses a simple JSONL structure.
Each example contains a Pashto grammar question or sentence together with an appropriate explanation or answer.
Example:
{
"messages": [
{
"role": "user",
"content": "ΩΎΩ Ψ―Ϋ Ψ¬Ω
ΩΩ Ϊ©Ϋ ΩΨΉΩ Ϊ©ΩΩ
Ψ―ΫΨ Β«Ψ§ΨΩ
Ψ― Ϊ©ΨͺΨ§Ψ¨ ΩΩΩΩ.Β»"
},
{
"role": "assistant",
"content": "ΩΎΩ Ψ―Ϋ Ψ¬Ω
ΩΩ Ϊ©Ϋ Β«ΩΩΩΩΒ» ΩΨΉΩ Ψ―Ϋ."
}
]
}The messages structure makes the dataset suitable for modern conversational language models and common SFT training pipelines.
π§ Grammar Coverage
The dataset provides examples covering several fundamental areas of Pashto grammar.
1. Sentence Structure
Examples help identify the basic structure of Pashto sentences, including:
- Subject
- Object
- Verb
- Sentence order
- Simple sentence construction
Pashto commonly follows an SOV (SubjectβObjectβVerb) pattern.
Example:
Ψ§ΨΩ Ψ― Ϊ©ΨͺΨ§Ψ¨ ΩΩΩΩ.
- Subject: Ψ§ΨΩ Ψ―
- Object: Ϊ©ΨͺΨ§Ψ¨
- Verb: ΩΩΩΩ
2. Verbs
The dataset includes examples involving Pashto verbs and their grammatical functions.
Examples include:
- Present tense
- Past tense
- Future tense
- Auxiliary usage
- Verb identification
- Verb agreement
- Simple verb constructions
3. Nouns
Examples demonstrate the identification and use of nouns in Pashto sentences.
Nouns may refer to:
- People
- Places
- Objects
- Animals
- Concepts
4. Pronouns
The dataset includes examples involving common Pashto pronouns and their grammatical roles.
Examples include:
- Ψ²Ω
- ΨͺΩ
- ΩΨΊΩ
- Ω ΩΪ
- ΨͺΨ§Ψ³Ω
- Ψ―ΩΫ
5. Adjectives
Examples demonstrate how adjectives describe nouns and how they function within Pashto sentences.
Example:
ΪΪ©ΩΫ Ϊ©ΩΨ±
Here:
- ΪΪ©ΩΫ β adjective
- Ϊ©ΩΨ± β noun
6. Adverbs
Examples demonstrate words that describe how, when, or where an action takes place.
7. Negation
The dataset includes examples of negative sentence construction.
Example:
Ψ²Ω ΪΩΩΩΪΩ ΨͺΩ ΩΩ ΪΩ .
The word ΩΩ is used to form the negative construction.
8. Questions
Examples include interrogative sentence structures.
Common question words include:
- Ϊ ΩΪ©Ψ
- Ϊ ΩΨ
- ΪΫΨ±ΨͺΩΨ
- Ϊ©ΩΩΨ
- ΩΩΫΨ
- Ϊ ΩΪ«ΩΨ
- Ϊ©ΩΩ Ψ
9. Tense
The dataset provides examples for recognizing different grammatical time references, including:
- Present
- Past
- Future
π§ͺ Intended Uses
This dataset can be used for:
π€ AI / NLP
- Pashto language model training
- Supervised fine-tuning (SFT)
- Instruction tuning
- Grammar evaluation
- Grammatical error analysis
- Pashto NLP research
- Low-resource language research
π Education
- Pashto language learning
- Grammar tutoring
- Classroom exercises
- Self-study
- Automated grammar explanation
π Evaluation
The dataset can also be used as a small evaluation set for testing whether a language model can:
- Identify grammatical components
- Recognize verb tense
- Understand sentence structure
- Explain basic grammatical rules
- Produce grammatically appropriate Pashto
βοΈ Training Recommendations
Because this is a small 100-example dataset, it should not normally be used as a standalone training corpus for a large language model.
Instead, it can be combined with larger Pashto datasets for:
- SFT
- Instruction tuning
- Grammar specialization
- Domain adaptation
- Evaluation
- Final-stage refinement
The dataset is particularly useful as a small, focused grammar component inside a larger Pashto training mixture.
π¬ Research Applications
Possible research applications include:
- Pashto grammar modeling
- Low-resource language modeling
- Pashto instruction following
- Grammatical classification
- Grammar-aware LLM evaluation
- Pashto educational AI
- Multilingual model adaptation
- Pashto tokenizer and language-model research
π§Ή Data Quality
The dataset is designed to remain:
- Focused on Pashto grammar
- Written in Pashto
- UTF-8 encoded
- Machine-learning friendly
- Simple to parse
- Suitable for conversational SFT formats
The examples are intended to emphasize grammatical understanding rather than general-domain knowledge.
β οΈ Limitations
This dataset contains only 100 examples and therefore should not be considered a comprehensive representation of the Pashto language.
It does not attempt to cover:
- All Pashto dialects
- Every grammatical construction
- Advanced morphology
- Complete verb paradigms
- Literary Pashto
- Specialized linguistic terminology
- All regional variations
For comprehensive Pashto grammar modeling, this dataset should be combined with larger and more diverse Pashto corpora.
π License
This dataset is released under the:
Creative Commons Attribution 4.0 International (CC BY 4.0)
You are free to:
- Share the dataset
- Adapt the dataset
- Use it for research
- Use it for machine-learning applications
- Use it in commercial projects
provided that appropriate attribution is given according to the CC BY 4.0 license.
π€ Author
Nasibullah Nassim (Nassim JP)
Pashto AI Research & Development
π―π΅ Japan
Hugging Face:
nassimjp
π Project
This dataset is part of the broader Pashto AI / iPashto.ai effort to develop open datasets and resources for Pashto language technology.
The broader goal is to improve the representation of Pashto language, grammar, reasoning, conversation, translation, and cultural knowledge in modern AI systems.
π Citation
If you use this dataset in research or a project, please cite:
@dataset{nassimjp_pashto_grammar_100_2026,
author = {Nasibullah Nassim},
title = {Pashto Grammar 100},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/nassimjp/Pashto-grammar-100},
license = {CC-BY-4.0}
}π€ Contributions
Contributions that improve Pashto language resources are welcome.
Possible future improvements include:
- Additional grammar examples
- More verb paradigms
- Morphological analysis
- Dialectal coverage
- Grammar correction examples
- Advanced syntax
- Linguistic annotations
- Native-speaker validation
- Larger evaluation sets
π¦π« Pashto AI
Pashto is a major language with millions of speakers, yet it remains significantly underrepresented in modern AI systems.
Small, focused datasets such as Pashto Grammar 100 can contribute to building better Pashto-capable language models by providing targeted linguistic knowledge that may be missing from large general-purpose corpora.
Ψ― ΩΎΪΨͺΩ ΩΩΎΨ§Ψ±ΩΨ ΩΨ±Ω ΩΎΨ§Ϊ©Ω Ψ§Ω Ϊ«ΩΌΩΨ±Ω Ϊ©Ψ±ΪΩ Ψ§Ψ±Ψ²ΪΨͺ ΩΨ±Ω. β€οΈπ¦π«
