nassimjp/Bilingual-SFT-Dataset
Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.
Update README.md
Upload bilingual.jsonl with huggingface_hub
Upload bilingual.jsonl with huggingface_hub
Update README.md
Create README.md
Upload bilingual.jsonl with huggingface_hub
initial commit
