CoolFace
Modelpublic

nassimjp/Ghanam-7B-Base-Pashto-v0.1

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes13downloads
Model Card

Ghanam-7B-Base-Pashto-v0.1

Model Description

Ghanam-7B-Base-Pashto-v0.1 is an experimental base model that extends the Mistral-7B-v0.1 vocabulary to include 47 Pashto/Arabic characters. This is the first stage of creating a Pashto-capable language model, where the tokenizer has been modified and the embedding layer has been expanded to support Pashto script.

Key Features

  • —Vocabulary Size: 32,010 tokens (original 32,000 + 47 Pashto characters)
  • —Quantization: 4-bit NF4 with double quantization for efficient inference
  • —Memory Efficient: ~4GB VRAM usage for inference
  • —Pashto Script Support: Now recognizes all Pashto-specific characters including ښ، ږ، څ، ډ، ړ، ځ، ګ، ۍ، ې

Pashto Characters Added

ا ب پ ت ث ج چ ح خ د ذ ر ز ژ س ش ص ض ط ظ ع غ ف ق ک ګ گ ل م ن ڼ و ه ی ډ ړ ځ څ ښ ږ ۍ ې ء آ أ ؤ ئ

⚠️ Important Note

This model does not yet understand Pashto language semantics. It can:

  • —✅ Tokenize Pashto text correctly
  • —✅ Generate Pashto characters (if prompted in Pashto)
  • —❌ Understand Pashto meaning or context
  • —❌ Respond meaningfully in Pashto

The model still generates English text as it hasn't been fine-tuned on Pashto data.

Model Details

Model Description

This model is the first step in creating a Pashto language model based on Mistral-7B. The vocabulary was expanded using the Stanford vocabulary expansion method (mean resizing), where new embeddings were initialized using the mean of existing embeddings to minimize disruption to the model's performance.

  • —Developed by: Nassim JP (Community Project)
  • —Model type: Causal Language Model (Decoder-only Transformer)
  • —Language(s): English (original) + Pashto script support (tokenization only)
  • —License: Apache 2.0
  • —Base Model: Mistral-7B-v0.1
  • —Quantization: 4-bit NF4 (BitsAndBytes)
  • —Vocabulary Expansion Method: Stanford mean resizing

Model Sources

Uses

Direct Use

You can use this model for:

  • —Pashto tokenization and text preprocessing
  • —Inference experiments with Pashto prompts (though responses will be in English)
  • —As a starting point for fine-tuning on Pashto datasets
  • —Research on vocabulary expansion for low-resource languages

Downstream Use

This model is intended to be fine-tuned on Pashto text data for:

  • —Pashto language modeling
  • —Pashto text generation
  • —Pashto translation tasks
  • —Pashto question answering

Out-of-Scope Use

  • —Do not use this model for production Pashto applications without fine-tuning
  • —Do not expect Pashto understanding or generation in Pashto
  • —Not suitable for tasks requiring Pashto semantic comprehension
  • —Not tested for bias, toxicity, or safety in Pashto context

Bias, Risks, and Limitations

Technical Limitations

  1. 1.No Pashto Understanding: The model has only expanded vocabulary but has not learned Pashto semantics
  2. 2.English-Only Knowledge: All pre-training knowledge remains in English
  3. 3.Potential Performance Degradation: Vocabulary expansion may slightly impact English performance
  4. 4.Untested on Pashto Tasks: No evaluation has been conducted on Pashto benchmarks

Recommendations

  • —Use this model only as a foundation for Pashto fine-tuning
  • —Validate carefully before using in any production environment
  • —Expect English responses when prompting in Pashto (until fine-tuned)
  • —Consider this a research checkpoint rather than a production model

How to Get Started with the Model

Loading the Model (4-bit)

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

# Model path (local or from Hub)
model_path = "nassimjp/Ghanam-7B-Base-Pashto-v0.1"

# 4-bit quantization config
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16
)

# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    quantization_config=bnb_config,
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_path)

# Set padding token (optional)
tokenizer.pad_token = tokenizer.eos_token

# Test tokenization
pashto_text = "پښتونخوا کې ډېر ښکلي ښارونه او ځنګلونه شته"
tokens = tokenizer.encode(pashto_text)
print(f"Tokens: {tokens}")
print(f"Decoded: {tokenizer.decode(tokens)}")

Generating Text

python
# Prompt in Pashto (model will respond in English)
prompt = "سلام، څنګه یاست؟"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=50,
        temperature=0.7,
        do_sample=True,
        repetition_penalty=1.2
    )

response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)

Inspecting New Tokens

python
# Check if Pashto characters are recognized
pashto_chars = ["ښ", "ږ", "څ", "ډ", "ړ", "ځ", "ګ"]
for char in pashto_chars:
    token_id = tokenizer.encode(char)[0]
    token_str = tokenizer.decode([token_id])
    print(f"Character: {char} -> Token ID: {token_id} -> Decoded: {token_str}")

Training Details

Training Data

No training data was used for this model. The vocabulary was expanded without any fine-tuning on Pashto text. The model retains only the original Mistral-7B pre-training knowledge.

Training Procedure

Vocabulary Expansion Method

The model underwent a two-stage process:

Stage 1: Tokenizer Modification

  • —Added 47 Pashto/Arabic characters to the original tokenizer
  • —Preserved all original tokens (no replacements)
  • —New vocabulary size: 32,010

Stage 2: Embedding Resizing

  • —Used Stanford vocabulary expansion method (mean resizing)
  • —New embeddings initialized with the mean of existing embeddings
  • —Both lm_head and embedding layers were resized
  • —No gradient updates or training performed
Training Hyperparameters
  • —Quantization: 4-bit NF4 with double quantization
  • —Compute dtype: bfloat16
  • —Method: Vocabulary expansion only (no fine-tuning)
  • —Embedding Initialization: Mean resizing

Evaluation

Testing Data, Factors & Metrics

Testing Data

Initial testing was performed using:

  • —Pashto sample texts to verify tokenization
  • —English prompts to ensure base functionality
  • —Character-level inspection to confirm tokenizer additions
Metrics
  • —Vocabulary Size: Successfully expanded from 32,000 to 32,010
  • —Tokenization Accuracy: All 47 characters are properly tokenized
  • —Generation Quality: English generation preserved
  • —Memory Usage: ~4GB VRAM for inference

Results

TestResultStatus
Pashto Tokenization✅ All 47 characters recognizedPASS
English Generation✅ Original performance preservedPASS
Pashto Understanding❌ No semantic understandingEXPECTED
Pashto Response❌ Generates EnglishEXPECTED
Model Loading (4-bit)✅ Successfully loadsPASS
Memory Efficiency✅ ~4GB VRAMPASS

Environmental Impact

Carbon emissions were minimal as this model:

  • —Did not require any training
  • —Only performed vocabulary expansion and quantization
  • —Used a single GPU for a few minutes
  • —Hardware Type: NVIDIA A100 (or similar)
  • —Hours used: < 1 hour
  • —Cloud Provider: N/A (local)
  • —Compute Region: N/A
  • —Carbon Emitted: Negligible

Technical Specifications

Model Architecture and Objective

  • —Architecture: Transformer decoder (Mistral-7B)
  • —Number of Layers: 32
  • —Hidden Size: 4096
  • —Attention Heads: 32
  • —Intermediate Size: 14336
  • —Vocabulary Size: 32,010 (expanded from 32,000)
  • —Positional Encoding: Rotary Position Embeddings (RoPE)

Compute Infrastructure

Hardware
  • —CPU for embedding expansion
  • —GPU for quantization and inference testing
Software
  • —Transformers: v4.31.0+
  • —BitsAndBytes: v0.41.0+
  • —PyTorch: v2.0.0+
  • —Accelerate: v0.20.0+

Model Card Authors

  • —Nassim JP - Vocabulary expansion and quantization implementation
  • —Original model: Mistral AI team

Model Card Contact

For questions or collaboration:

Acknowledgments

  • —Mistral AI for the base model
  • —Stanford NLP for the vocabulary expansion method
  • —Hugging Face for the transformers library
  • —BitsAndBytes team for 4-bit quantization

Glossary

  • —Vocabulary Expansion: Adding new tokens to a model's tokenizer
  • —Mean Resizing: Initializing new embeddings with the mean of existing embeddings
  • —4-bit NF4: 4-bit Normal Float quantization method
  • —Fine-tuning: Training a pre-trained model on domain-specific data

Next Steps

To make this model truly Pashto-capable, the next steps include:

  1. 1.Collect Pashto corpus (books, articles, web text)
  2. 2.Pre-train on Pashto data (continued pre-training)
  3. 3.Fine-tune for specific tasks (translation, QA, generation)
  4. 4.Evaluate on Pashto benchmarks
  5. 5.Deploy in production environments

Citation

If you use this model or the vocabulary expansion method, please cite:

bibtex
@misc{ghanam2024pashto,
  author = {Nassim JP},
  title = {Ghanam-7B-Base-Pashto: Vocabulary Expansion for Pashto Language},
  year = {2024},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/nassimjp/Ghanam-7B-Base-Pashto-v0.1}}
}

Original Mistral-7B Citation

bibtex
@article{jiang2023mistral,
  title={Mistral 7B},
  author={Jiang, Albert Q and others},
  journal={arXiv preprint arXiv:2310.06825},
  year={2023}
}

Additional Information

Usage Tips

  1. 1.Tokenization: The tokenizer now correctly handles all Pashto characters
  2. 2.Memory: 4-bit quantization allows running on consumer GPUs with 6-8GB VRAM
  3. 3.Generation: Use do_sample=True and temperature=0.7 for diverse outputs
  4. 4.Pad Token: If using batch generation, set pad_token=tokenizer.eos_token

Common Issues

IssueSolution
Tokenizer missing charactersEnsure you're using the correct tokenizer from this repo
Memory errorsReduce max_new_tokens or use CPU offloading
Poor generation qualityThis is expected - the model needs fine-tuning
Pashto text displayed as [UNK]Tokenizer hasn't loaded properly - reload from the repo

Model Status: 🔬 Research Phase - Pre-fine-tuning checkpoint

Last Updated: June 2026