CoolFace
Datasetpublic

Jamalianpour/persian-wikipedia-instruct

Persian Wikipedia Instruct A Persian-language instruction-tuning dataset of ~120,000 samples, generated from the Persian Wikipedia (fawiki) article dump. Each article was cleaned to Markdown, chunked by section, and passed to a locally-run Gemma4 model that produced grounded instruction/response pairs across five task types. The result is ready for supervised fine-tuning (SFT) of Persian LLMs. Heads-up: this is synthetic data. The instructions and answers were written by an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Jamalianpour/persian-wikipedia-instruct.

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
0likes88downloads
4 commits on main
d45a4403mo ago

Update README.md

Jamalianpour
173cf673mo ago

Upload folder using huggingface_hub

Jamalianpour
6acef0c3mo ago

Upload folder using huggingface_hub

Jamalianpour
508f34e3mo ago

initial commit

Jamalianpour