datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian_bbh
Persian BBH
This is BIG-bench Hard dataset translated to Persian using GPT-4o-mini.
We use 19 out of 23 original tasks in BIG-bench Hard.
PersianSyntheticQA
Persian Synthetic QA Dataset
Persian Synthetic QA is a dataset containing 100,000 synthetic questions and answers in Persian, generated using GPT-4o. The dataset is structured as conversations between a user and an assistant, with 2,000 records for each of the 50 different topics. Each conversation consists of messages with two distinct roles: "user" messages containing questions in Persian, and "assistant" messages containing the corresponding answers. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ParsBench/PersianSyntheticQA.iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.persian-poetry-metersPersian poems with their corresponding meters, from ganjoor.net.
iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.PersianSyntheticQA
Persian Synthetic QA Dataset
Persian Synthetic QA is a dataset containing 100,000 synthetic questions and answers in Persian, generated using GPT-4o. The dataset is structured as conversations between a user and an assistant, with 2,000 records for each of the 50 different topics. Each conversation consists of messages with two distinct roles: "user" messages containing questions in Persian, and "assistant" messages containing the corresponding answers. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/PersianSyntheticQA.persian-poetics-kb
Persian Poetics Knowledge Base — پایگاه دانش شعر و رپ فارسی
مجموعهٔ بازیابی برای ساخت و ارزیابی شعر/رپ فارسی بهمثابهٔ پرسشوپاسخِ
محدودیتدار و مبتنی بر شواهد (PersianPoet-RAG v2). هر واحد هم حاشیهنویسی
معنایی (معنا، دامنه، تصویر) دارد و هم آوایی (واجها، ساخت هجا، تکیه،
کلید قافیهٔ سختگیرانه/آسانگیر، زنجیرهٔ واکهها) — چیزی که بازیاب را قادر
میکند وزن و قافیه را قبل از فراخوانی مدل زبانی تأمین کند.
چه چیزهایی داخل این دیتاست است؟ (What's inside)
کانفیگ… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/persian-poetics-kb.Maux-Persian-SFT-30k
Maux-Persian-SFT-30k
Dataset Description
This dataset contains 30,000 high-quality Persian (Farsi) conversations for supervised fine-tuning (SFT) of conversational AI models. The dataset combines multiple sources to provide diverse, natural Persian conversations covering various topics and interaction patterns.
Dataset Structure
Each entry contains:
messages: List of conversation messages with role (user/assistant/system) and content
source: Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/Maux-Persian-SFT-30k.clinical-persian-qa-ii
Clinical Question Answering Dataset II (Farsi)
This dataset contains more than 211k questions and more than 700k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset is NOT a part of Clinical Question Answering I dataset and is a whole… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-ii.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.persian-wikipedia-instruct
Persian Wikipedia Instruct
A Persian-language instruction-tuning dataset of ~120,000 samples, generated from the
Persian Wikipedia (fawiki) article dump. Each article was cleaned to Markdown, chunked by
section, and passed to a locally-run Gemma4 model that produced grounded instruction/response
pairs across five task types. The result is ready for supervised fine-tuning (SFT) of
Persian LLMs.
Heads-up: this is synthetic data. The instructions and answers were written by an
LLM… See the full description on the dataset page: https://huggingface.co/datasets/Jamalianpour/persian-wikipedia-instruct.PersianMedQA
PersianMedQA
PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
A large-scale, expert-validated multiple-choice question set covering 23 medical specialties, collected from 14 years (2011–2024) of Iranian national residency and pre-residency board examinations administered by Sanjeshp (the Medical Education Assessment Center, under the Iranian Ministry of Health).
Total items: 20,785
Train 14,549 · Validation 1,000… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA.persian-gk-cleanedThis is a cleaned and validated version of the original mshojaei77/persian-gk dataset.
The purpose of this version is to ensure robust compatibility with modern fine-tuning workflows that rely on strict chat templates (e.g., tokenizer.apply_chat_template). The cleaning process resolves structural errors in the original dataset that could cause TemplateError or other silent failures during training with models like Gemma 3N, Llama 3, and others.
Cleaning and Validation Process… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-gk-cleaned.crossword-puzzle-persian-cheat
Crossword Puzzle Cheat Dataset (Persian)
This dataset consists of 30157 pairs of questions and answers.
Dataset Description
The reference for this dataset is jadvalyab.ir website.
Usage
Huggingface datasets library:
from datasets import load_dataset
dataset = load_dataset('PerSets/crossword-puzzle-persian-cheat')
License
CC0-v1.0
persian-nlu
Persian NLU
Dataset Summary
The Persian NLU Benchmark is a curated collection of existing Persian datasets designed to evaluate Natural Language Understanding (NLU) capabilities across a diverse range of tasks. It provides a unified benchmark suite to assess different cognitive aspects of large language models (LLMs) in Persian.
This benchmark includes the following tasks and datasets:
Text Classification:
Synthetic Persian Tone
SID
Natural Language Inference… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-nlu.clinical-persian-qa-i
Clinical Question Answering Dataset I (Farsi)
This dataset contains approximately 50k questions and around 60k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset is NOT a part of Clinical Question Answering II dataset and is a complete… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-i.drhast-persian-medical-QA
Medical Question and Answer Dataset
Crawled from https://drhast.com.
Dataset Structure
{
"title": "پایین امدن دیابت بارداری و ادامه دادن قرص گلوکوفاژ؟",
"url": "https://drhast.com/q/yLvD6",
"text": "سلام من هفته ۳۳ بارداری هستم هفته ۳۱ ازمایش دادم تشخیص دیابت دادن ناشتا۹۳ یکساعت بعد خوردن محلول ۲۰۷ و دوساعت بعد۱۷۲ متخصص داخلی روزی دوتا گلوکوفاژ و یه سری محدودیتا تغذیه مشخص کردن دوکیلو وزن کم کردم و امروز مجدد ازمایش دادم قند ناشتا ۸۴ و دوساعت بعد… See the full description on the dataset page: https://huggingface.co/datasets/AliMoameri/drhast-persian-medical-QA.persian-med-qa
🏥 Persian Medical Question Answering Dataset
Dataset Summary
The Persian Medical QA Dataset is a high-quality, expert-curated collection of question–answer (QA) pairs in Persian (Farsi), designed for developing and evaluating natural language processing (NLP) models for medical question answering. All answers are grounded in reliable medical resources, including standard medical textbooks (e.g., Harrison’s Principles of Internal Medicine) and authoritative medical… See the full description on the dataset page: https://huggingface.co/datasets/aictsharif/persian-med-qa.Persian_Riddles
Persian Riddles (چیستانهای فارسی)
A curated collection of 1,065 Persian-language riddles, each paired with its answer. The set spans classic Persian riddles (چیستان), wordplay and letter puzzles, lateral-thinking brain-teasers, and knowledge-based trick questions gathered from a range of Persian sources.
Duplicates were removed during compilation: entries were normalized for Persian/Arabic character variants (ی/ي, ک/ك), zero-width joiners, diacritics, and punctuation, and a… See the full description on the dataset page: https://huggingface.co/datasets/yoyo-research-group/Persian_Riddles.persian-legal-kb
Persian Legal Knowledge Base — پایگاه دانش حقوقی فارسی
A provenance-first Persian legal corpus built for retrieval-augmented and
knowledge-graph (KAG) systems, not for fine-tuning knowledge into weights.
Every record answers four questions a legal AI must never get wrong:
is this law or just a bill? · is it still in force? · which article? · who said so?
Why this dataset exists
Iranian law moved in the two weeks before this dataset was built:
قانون ملی توسعه هوش… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/persian-legal-kb.parsinlu_reading_comprehensionA Persian reading comprehenion task (generating an answer, given a question and a context paragraph).
The questions are mined using Google auto-complete, their answers and the corresponding evidence documents are manually annotated by native speakers.persian-gk
Dataset Card for persian-gk (Persian General Knowledge)
Dataset Summary
persian-gk is a cleaned and structured collection of Persian (Farsi) conversation pairs covering a wide range of general-knowledge topics. Each conversation is formatted in ChatML style with explicit system, user, and assistant roles, enabling straightforward use for both instruction-tuning and chat-style language-model training.
Language: Persian (fa)
Size: 5 897 conversations, 2–8 turns each (≈… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-gk.ZharfaTech-Open-Platypus-Persian-Farsi
Persian Open-Platypus
About ZharfaTech
ZharfaTech is a pioneer in developing Language Learning Models (LLMs) tailored for the Persian language, aiming to empower over 100 million Persian speakers worldwide. Our mission encompasses bridging the digital divide in LLM-related services like content generation, customer relationship systems, and more, with a dual approach of fostering open-source collaboration and delivering high-value, specialized closed-source solutions.… See the full description on the dataset page: https://huggingface.co/datasets/ZharfaTech/ZharfaTech-Open-Platypus-Persian-Farsi.persian-natural-fluently
Persian scientific dataset
I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology.
The content of the dataset has been generated by : Human, Grok3, DeepSeek R1.
License
This dataset is licensed under apache-2.0.
mauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/mauxi-COT-Persian.Persian-Thinking
Persian-Thinking
Persian-Thinking is a small Persian-language reasoning/thinking dataset created by sampling and translating a subset of SmolTalk2.
Dataset Details
1,000 samples (999 after processing) drawn from the smoltalk_systemchats_Qwen3_32B_think subset of SmolTalk2, part of its SFT split.
That source subset consists of system-chat conversations generated with Qwen3-32B in thinking mode, meaning each assistant response includes an explicit reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/Persian-Thinking.PersianMHQAThe PersianMHQA Dataset is the first open-domain, multi-hop question answering dataset, containing 7,000 questions generated using Persian Wikipedia texts. We have currently made available 500 samples from the dataset records to introduce this dataset. Soon, after the publication of the paper, the entire dataset will be accessible.
Contact Us
** Mobina Taji Email 📧**: mtaji3405@gmail.com
ErfanClone-persian-alpaca
Persian Alpaca
This dataset is directly cloned from this one. All the credits should go to iamshnoo for translating alpaca to persian.
Translated from yahma/alpaca-cleaned using NLLB-1.3B
Persian-MuSR
Persian MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning on Persian Language
This is the Persian-translated version (using GPT-4o) of the original dataset MuSR.
Acknowledgments
Special thanks to AvalAI for sponsoring this project through their AvalAward program
This dataset was made possible by AvalAI's generous support and commitment to advancing Persian language AI research
