CoolFace
Datasetpublic

nassimjp/Bilingual-SFT-Dataset

Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes67downloads
7 commits on main
5d0bcb21mo ago

Update README.md

nassimjp
5cffd961mo ago

Upload bilingual.jsonl with huggingface_hub

nassimjp
71f24e91mo ago

Upload bilingual.jsonl with huggingface_hub

nassimjp
4de11261mo ago

Update README.md

nassimjp
fc8c6e91mo ago

Create README.md

nassimjp
5eab6031mo ago

Upload bilingual.jsonl with huggingface_hub

nassimjp
3f51f6e1mo ago

initial commit

nassimjp