datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-abusive-words
Persian Abusive Words Dataset
This is a labeled dataset of Persian Abusive Words, originally sourced from Persian Abusive Words GitHub repository. The dataset has been split by the contributor into two subsets: train and test.
This dataset can be used for developing systems to detect and filter offensive or abusive language in various contexts. It is particularly useful for identifying inappropriate words and managing content moderation in applications where Persian language… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaFzp/persian-abusive-words.peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.persian-poetry-metersPersian poems with their corresponding meters, from ganjoor.net.
persian-poetry-qa
Persian Poetry Dataset
Dataset Description
Overview
This dataset contains a collection of Persian poems structured in a question-answering format. The dataset is derived from various Persian poets and their poems, providing a rich source for exploring Persian poetry in a structured manner suitable for machine learning applications, especially in natural language processing tasks like question answering.
Data Collection
Data Collection Source: The… See the full description on the dataset page: https://huggingface.co/datasets/kakooch/persian-poetry-qa.persian-handwriting-ocr
Persian Handwriting OCR Dataset
Dataset Summary
A standardized dataset of Persian (Farsi) handwritten pages with word-level
bounding-box annotations and transcriptions. The dataset is page-level:
each sample is a full page scan; annotations are one row per word bbox on
that page. This is the most flexible form -- users can train page-level OCR,
word detection (DBNet/PaddleOCR), or derive word/line crops as needed.
Pages: 1115 scanned pages (canonical IDs… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-handwriting-ocr.Persian_poemPersian-Food-Sentiment
Persian Food Sentiment Dataset
This data is orinally from https://hooshvare.github.io/docs/datasets/sa.
BibTeX Citation
If you use this dataset, please cite following paper:
@article{ParsBERT,
title={ParsBERT: Transformer-based Model for Persian Language Understanding},
author={Mehrdad Farahani, Mohammad Gharachorloo, Marzieh Farahani, Mohammad Manthouri},
journal={ArXiv},
year={2020},
volume={abs/2005.12515}
}
English-Persian-Parallel-Dataset
English-Persian Parallel Dataset
This repository provides access to a high-quality parallel dataset for English-to-Persian translation. The dataset has been curated for research purposes and is suitable for training and evaluating Neural Machine Translation (NMT) models.
Download Link
You can download the dataset using the following link:
Download English-Persian Parallel Dataset
Description
The dataset contains aligned sentence pairs in English and Persian… See the full description on the dataset page: https://huggingface.co/datasets/shenasa/English-Persian-Parallel-Dataset.Persian_sentimentPersian-OCR-230k.
├── Images/
├── train.csv (184k)
└── test.csv (46k)
crossword-puzzle-persian-cheat
Crossword Puzzle Cheat Dataset (Persian)
This dataset consists of 30157 pairs of questions and answers.
Dataset Description
The reference for this dataset is jadvalyab.ir website.
Usage
Huggingface datasets library:
from datasets import load_dataset
dataset = load_dataset('PerSets/crossword-puzzle-persian-cheat')
License
CC0-v1.0
persian-gender-by-name
Persian Gender Detection by Name
A comprehensive dataset for determining gender based on Persian names, enriched with English representations.
Overview
The Persian Gender Detection by Name dataset is the largest of its kind, comprising approximately 27,000 entries. Each entry includes a Persian name, its corresponding gender, and the English transliteration. This dataset is designed to facilitate accurate gender detection and enhance searchability through multiple name… See the full description on the dataset page: https://huggingface.co/datasets/farbodbij/persian-gender-by-name.PersianMedQA
PersianMedQA
PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
A large-scale, expert-validated multiple-choice question set covering 23 medical specialties, collected from 14 years (2011–2024) of Iranian national residency and pre-residency board examinations administered by Sanjeshp (the Medical Education Assessment Center, under the Iranian Ministry of Health).
Total items: 20,785
Train 14,549 · Validation 1,000… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA.Collection-of-drug-names-in-Persian
Dataset Details
A collection of drug names in Persian
Language(s) (NLP): persian
PersianTwitterDataset-SentimentAnalysis
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains more than 3300 Persian tweets, crawled from X.com
Each tweet is assigned a label, which is a number between 0 to 4.
Label 0 indicates the sentiment of Happiness and Joy.
Label 1 indicates the sentiment of Sadness.
Label 2 indicates the sentiment of Anger and… See the full description on the dataset page: https://huggingface.co/datasets/moali-mkh-2000/PersianTwitterDataset-SentimentAnalysis.ncbi-persian
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Amir13/ncbi-persian.persian-med-qa
🏥 Persian Medical Question Answering Dataset
Dataset Summary
The Persian Medical QA Dataset is a high-quality, expert-curated collection of question–answer (QA) pairs in Persian (Farsi), designed for developing and evaluating natural language processing (NLP) models for medical question answering. All answers are grounded in reliable medical resources, including standard medical textbooks (e.g., Harrison’s Principles of Internal Medicine) and authoritative medical… See the full description on the dataset page: https://huggingface.co/datasets/aictsharif/persian-med-qa.ontonotes5-persian
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Amir13/ontonotes5-persian.RAG_vs_FineTuning_Comparison_Persian_V2chatgpt-prompts-Persianconll2003-persian
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Amir13/conll2003-persian.mauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/mauxi-COT-Persian.RAG_vs_FineTuning_Comparison_Persian_V1SynPerForm-synthetic-persian-formality-pairs
SynPerForm
SynPerForm is a paired Persian dataset for formality style
transfer. Each informal text is paired with a freely written formal rewrite
that preserves its meaning without requiring lexical or structural
equivalence. The formal rewrites were generated with OpenAI GPT-5.6 Luna.
Columns
Informal: the original informal Persian text.
Formal: a free formal rewrite that preserves the original meaning.
Intended uses
Formality style transfer:… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/SynPerForm-synthetic-persian-formality-pairs.conversational-persian-subtitles
Conversational Persian Subtitles
Dataset name: Conversational Persian SubtitlesCollaboration: Maral Zarvani & Milad Ghashangi AgdamLicense: CC BY 4.0Hugging Face Repo: https://huggingface.co/datasets/Maral/conversational-persian-subtitles
1. Dataset Description
This dataset contains cleaned Persian subtitle lines from a wide variety of Korean TV series and films, each line reflecting informal, conversational dialogue. All markup (square brackets, timecodes,etc.) has… See the full description on the dataset page: https://huggingface.co/datasets/Maral/conversational-persian-subtitles.persian-alpaca-deep-clean
Persian Alpaca Deep Clean
Overview
The Persian Alpaca Dataset is a collection of finely cleaned Persian language records derived from various sources, primarily the Bactrian, PN-Summary (summarization), and PEYMA (Named Entity Recognition) datasets. The dataset comprises approximately 68,279 records after rigorous cleaning processes, including character normalization, removal of Arabic letters, elimination of sentences with high word repetition, removal of words with… See the full description on the dataset page: https://huggingface.co/datasets/myrkur/persian-alpaca-deep-clean.alpaca_persian_adabimauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/mauxi-COT-Persian.Persian_Cookingpersian-ai-generated-text
📝 Persian AI-Generated Text Dataset
A large-scale collection of 10,546 AI-generated Persian (Farsi) texts produced by 70 different large language models across diverse topics and writing styles. This dataset is designed to support research in AI-generated text detection for the Persian language.
Dataset Summary
Attribute
Value
Language
Persian (Farsi)
Total Samples
10,546
Unique Models
70
API Providers
9 (OpenRouter, NVIDIA, Free Endpoint, HF… See the full description on the dataset page: https://huggingface.co/datasets/ehsantorabi/persian-ai-generated-text.
