datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.vibevoice-quran_persian-single-speakerpersian-ocr-community-dataset-argilla
Persian OCR community dataset - Argilla view
Lightweight two-column view for Argilla. image is an HF-hosted asset URL and label is JSON containing spatial OCR objects.
persian-license-plate-v1
Dataset is downloaded from here which was provided at Amirkabir University of Technology.
The dataset is labeled by the authors.
Experimental results show that the fine-tuned model works well in Persian License Plate.
Usage
You can download the dataset easily using HF datasets package in Python:
!pip install datasets
from datasets import load_dataset
dataset = load_dataset("hezarai/persian-license-plate-v1", split="train") # Other splits: validation, test
print(dataset[0])
vibevoice-gptinformal_persian-single-speakerpersian-handwriting-pages-3.69m
Persian Handwriting Pages 3.69M
3,690,000 deterministic, densely composed Persian handwriting pages.
This expansion uses new random seeds and is complementary to
Reza2kn/persian-handwriting-pages-369k,
not a repetition of its rendered pages.
The public viewer intentionally exposes exactly two columns: image and label.
Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification.
Source handwriting
Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.persian-ocr-community-datasetPersian-Farsi-Speech
Persian (Farsi) TTS Dataset
🗂️ Dataset Description
This dataset is a Persian (Farsi) text-to-speech (TTS) corpus built by concatenating and denoising multiple existing Farsi datasets.It is intended for training and evaluation of speech synthesis (TTS) models in Persian.
Since the basic datasets were contaminated with unintelligible audio, I used dnsmos to keep only clean audio (mos_ovr >= 3.0, same value as for the Emilia dataset).
The dataset contains two main… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/Persian-Farsi-Speech.persian-printed-ocr-3.5m
Persian Printed OCR 3.5M
A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public
datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733
rejected rows are excluded. The viewer exposes exactly image and label.
Sources
AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0)
hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.persian-handwritten-digits
Persian Handwritten Digits (Farsi)
80,000 grayscale images of handwritten Persian (Farsi) digits — ۰۱۲۳۴۵۶۷۸۹ —
organized as an ImageFolder dataset with 10 classes (0–9), 8,000 images per class.
Each image is a 28×28 grayscale PNG of a single digit.
Classes
Class
Count
0 (۰)
8,000
1 (۱)
8,000
2 (۲)
8,000
3 (۳)
8,000
4 (۴)
8,000
5 (۵)
8,000
6 (۶)
8,000
7 (۷)
8,000
8 (۸)
8,000
9 (۹)
8,000
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Mehdinmz/persian-handwritten-digits.persian-handwriting-pages-369k
Persian Handwriting Pages 369K
Full-page Persian handwriting compositions on scanned paper backgrounds.
Each row deliberately has only two fields:
image: the composed full-page image
label: its complete line-separated Persian transcription, ordered from top to bottom
The pages are composed from labeled real handwriting crops with page-level ink normalization,
controlled RTL layout variation, collision prevention, and exact transcription provenance.
The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.persianvox_2_rawpersian_bbh
Persian BBH
This is BIG-bench Hard dataset translated to Persian using GPT-4o-mini.
We use 19 out of 23 original tasks in BIG-bench Hard.
PersianSyntheticQA
Persian Synthetic QA Dataset
Persian Synthetic QA is a dataset containing 100,000 synthetic questions and answers in Persian, generated using GPT-4o. The dataset is structured as conversations between a user and an assistant, with 2,000 records for each of the 50 different topics. Each conversation consists of messages with two distinct roles: "user" messages containing questions in Persian, and "assistant" messages containing the corresponding answers. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ParsBench/PersianSyntheticQA.Persian_sentimentpersianvox_2_audioiran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.Finglish-To-Persian-Dataset-Large
Finglish to Persian Large Dataset
A massive-scale parallel corpus containing over 9.8 million sentence pairs for Finglish (Latin-script Persian) to Persian script transliteration. This dataset provides a robust foundation for training and fine-tuning seq2seq models, normalizing user-generated text, and enhancing Persian input methods.
What is Finglish?
Finglish (also known as Pinglish) is the practice of writing Persian using the Latin alphabet. Because there is… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Finglish-To-Persian-Dataset-Large.visualears-persian-asr-16k
🗂️ visualears-persian-asr-16k
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Main public Persian ASR audio/text dataset: 3.93M 16 kHz rows.
مجموعهدادهٔ اصلی و عمومی شنوا برای آموزش بازشناسی گفتار فارسی؛ شامل صوت ۱۶ کیلوهرتز، متن و فرادادهٔ منشأ در مقیاس چندمیلیونی.
🧩 Role
flagship training corpus
پیکرهٔ اصلی آموزش
📦 Snapshot
176 files; approximately 512.17 GB
176… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-persian-asr-16k.persian-ocr-bench-submitted10-bbox-crops
Persian OCR benchmark — selected submitted bbox crops
This dataset contains the non-empty OCR bboxes from the ten explicitly selected
submitted pages in persian_ocr_bench_bbox_review.
Each row is one PNG crop. gold_text is the current editable OCR content from
the live Argilla bbox field (content_text). Geometry is stored both as source
page pixels and as percentages of the source page. The original record ID,
external ID, bbox ID, source URL, and SHA-256 hashes are included for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-bbox-crops.persian-abusive-words
Persian Abusive Words Dataset
This is a labeled dataset of Persian Abusive Words, originally sourced from Persian Abusive Words GitHub repository. The dataset has been split by the contributor into two subsets: train and test.
This dataset can be used for developing systems to detect and filter offensive or abusive language in various contexts. It is particularly useful for identifying inappropriate words and managing content moderation in applications where Persian language… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaFzp/persian-abusive-words.Common-Voice-Speech-26.0-Persian-Clean
Persian Common Voice Clean Dataset
This dataset is a cleaned and prepared subset of the Persian (فارسی - fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12.
The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences and come from 3,791 speakers.
The original Persian Common Voice release contains… See the full description on the dataset page: https://huggingface.co/datasets/pymmdrza/Common-Voice-Speech-26.0-Persian-Clean.filimo-persian-asrThis dataset consists of about 400 hours of audio extracted from various Filimo videos in the Persian language.
Note: This dataset contains raw, unvalidated transcriptions. Users are advised to:
1. Perform their own quality assessment
2. Create their own train/validation/test splits based on their specific needs
3. Validate a subset of the data if needed for their use casePersian-Text-SentimentDataset Classes
negetive :0
positive :1
bookroom-persian-book-covers-and-titlesyoutube-persian-asrThis dataset consists of over 385 hours of audio extracted from various YouTube videos in the Persian language.
Note: This dataset contains raw, unvalidated transcriptions. Users are advised to:
1. Perform their own quality assessment
2. Create their own train/validation/test splits based on their specific needs
3. Validate a subset of the data if needed for their use casepersian-accents-benchmark
Persian Accents Benchmark
Dataset Summary
A benchmark for Persian automatic speech recognition (ASR): 279 short
utterances of informal Persian (Farsi) dialect speech across 16 regional accents,
released as a fixed evaluation set. Total audio duration is approximately 4.4
hours. The primary label is the transcription; each utterance also carries an
accent label (usable for accent classification as a secondary task) and an
emotion label as auxiliary metadata.
This… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-accents-benchmark.Persian-Food-Sentiment
Persian Food Sentiment Dataset
This data is orinally from https://hooshvare.github.io/docs/datasets/sa.
BibTeX Citation
If you use this dataset, please cite following paper:
@article{ParsBERT,
title={ParsBERT: Transformer-based Model for Persian Language Understanding},
author={Mehrdad Farahani, Mohammad Gharachorloo, Marzieh Farahani, Mohammad Manthouri},
journal={ArXiv},
year={2020},
volume={abs/2005.12515}
}
Persian_Common_Voice_17_0Tabaghe16_dataset_persian
