datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Persian-Farsi-Speech
Persian (Farsi) TTS Dataset
🗂️ Dataset Description
This dataset is a Persian (Farsi) text-to-speech (TTS) corpus built by concatenating and denoising multiple existing Farsi datasets.It is intended for training and evaluation of speech synthesis (TTS) models in Persian.
Since the basic datasets were contaminated with unintelligible audio, I used dnsmos to keep only clean audio (mos_ovr >= 3.0, same value as for the Emilia dataset).
The dataset contains two main… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/Persian-Farsi-Speech.persian-accents-benchmark
Persian Accents Benchmark
Dataset Summary
A benchmark for Persian automatic speech recognition (ASR): 279 short
utterances of informal Persian (Farsi) dialect speech across 16 regional accents,
released as a fixed evaluation set. Total audio duration is approximately 4.4
hours. The primary label is the transcription; each utterance also carries an
accent label (usable for accent classification as a secondary task) and an
emotion label as auxiliary metadata.
This… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-accents-benchmark.Persian-ASR-BenchmarkThis dataset consists of 3 hours of 16kHz audio collected from diverse environments to better represent real-world scenarios. The recordings were sourced from audiobooks, YouTube, and other public sources, ensuring a wide variety of speech styles and acoustic conditions.
One key advantage of this dataset is that it was collected from recent sources within the last few months, ensuring no overlap with training data and fairness for evaluating other STT models.
To enable a robust and fair… See the full description on the dataset page: https://huggingface.co/datasets/C1Tech/Persian-ASR-Benchmark.persian-tarjoman-koochik-16k
Persian Tarjoman — Koochik-labeled (16k mono FLAC, NeMo-tarred, streamable)
74,808 clips / 182.9 hours of clean long-form Persian narration (Tarjoman magazine articles).
Audio: from farsi-asr/PerSets-tarjoman-chunked, resampled to 16kHz mono.
Labels: transcribed with Reza2kn/Shenava-Koochik-v1.0 (CTC head) — the source's own Speechmatics transcripts were empty.
Format: NeMo-tarred / streamable — audio/shard_XXXXX.tar (19 shards) + manifests/train.jsonl (audio_filepath =… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-tarjoman-koochik-16k.persian-voice-v1
🗣️ Common Voice 17 — Persian (Spelling-Corrected Edition)
This is a refined version of the Persian subset of Mozilla's Common Voice 17 dataset, specially curated to enhance the performance of ASR (Automatic Speech Recognition) systems in Persian.
🛠️ Why this matters
The original dataset contained a significant number of spelling inconsistencies and typographical errors, which negatively impacted transcription accuracy and model alignment.
✨ What’s improved… See the full description on the dataset page: https://huggingface.co/datasets/vhdm/persian-voice-v1.persian-elderly-asr
Final gathered Persian elderly speech
Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history.
80 paired recordings from four speaker folders. Reference transcripts were aligned with a historical Persian Wav2Vec2-base checkpoint, then cut at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-elderly-asr.persianvox_all
PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data
PersianVox-All is the larger, less-filtered counterpart of PersianVox: a multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It contains every utterance that passed language and speech-quality (MOS) filtering, without the additional dual-ASR transcript-agreement filtering applied to the main PersianVox release. It is therefore substantially… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox_all.persian-accent-dataset
Persian Accent Speech Dataset
This dataset is an initial collection of Persian speech intended for accent
research and automatic-speech-recognition experiments. The first release
contains an esfahani split assembled from 22 YouTube videos selected as
Esfahani-accent source material.
Dataset contents
The current release contains:
479 audio clips
22 source videos
One split: esfahani
Two columns: audio and label
Audio normalized to mono, 16 kHz, 16-bit PCM WAV
Clip… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-accent-dataset.persianvox
PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data
PersianVox is a 2,400-hour, multi-speaker Persian (Farsi) speech corpus automatically mined from in-the-wild unlabeled data. It is, to date, the largest open-source speech resource for Persian, built to support zero-shot text-to-speech (TTS) research and other speech tasks in low-resource-language settings.
Dataset Summary
Advancement of zero-shot text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/persianvox.persian-words
Persian Words
This is a dataset of approximately 5K words, read aloud by a variety of native speakers. The dataset has been directly redistributed from this URL.
It can be used as a valuable resource for evaluating/training ASR engines or speech synthesis engines.
P.S.: I'm not the original creator of this dataset, for crediting or ownership change you can contact
persian-asr-text-2.69M-deduped
🗂️ persian-asr-text-2.69M-deduped
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Deduplicated Persian ASR text dataset used by the training stack.
پیکرهٔ متنی فارسیِ حذفتکرارشده برای ساخت واژگان، مدلسازی زبانی و پشتیبانی از آموزش ASR.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
4 files; approximately 109.64 MB
4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.Creativeguy-persian-asr
persian ASR Dataset
Audio resampled to 16000 Hz. Transcripts in Persian.
Splits
Split
Samples
Total Duration (H:MM:SS)
Volume
train
13165
100:37:38
2.712 GB
test
5643
43:08:16
1.163 GB
total
18808
143:45:55
3.874 GB
persian-asr-dataset
Ganjoor Persian Speech Dataset
مجموعهای فارسی برای بازشناسی گفتار که از محتوای صوتی و متنهای متناظر وبسایت گنجور گردآوری و به قالب Hugging Face تبدیل شده است. این مجموعه شامل گفتار سالمندان نیست و از دیتاست اختصاصی سالمندان پروژه مستقل است.
This is a Persian automatic speech recognition dataset derived from aligned audio and text crawled from Ganjoor. It is separate from the project's elderly-speech dataset.
Dataset structure
Split: train
Samples: 5,036
Audio:… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-asr-dataset.PersianAudiobook
PersianAudiobook
PersianAudiobook is a collection of short Persian audiobook speech clips paired with conservatively refined pseudo-transcriptions. The initial release contains 39,454 accepted examples representing 218.564 hours of mono, 16 kHz speech. It is intended for speech-recognition research, audiobook-domain language modeling, speech representation learning, and carefully reviewed text-to-speech research.
Raw data was collected from IranSeda audiobooks.
The… See the full description on the dataset page: https://huggingface.co/datasets/Pooya-Fallah/PersianAudiobook.costumer-service-persian-1PersianVox_HB
PersianVox_HB
PersianVox_HB is a high-quality, multispeaker Persian speech dataset derived from audio recordings of the Holy Bible. The dataset is sourced from bible.com (PCB=49.85 hours, TPV=71.4 hours), bible.is (NMV=24.17 hours), and wordproject.org (PHB=81.92 hours).
📚 Dataset Summary
Language: Persian (Farsi)
Speakers: Multiple speakers
Total Duration: 227.34 hours
Recording Sources:
bible.com (PCB: 49.85 hours, TPV: 71.4 hours)
bible.is (NMV: 24.17 hours)… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/PersianVox_HB.PersianVox_NM
PersianVox_NM
PersianVox_NM is a high-quality Persian speech dataset derived from article readings by a single female speaker. This subset is sourced from Nasle Mana Magazine, a publication dedicated to producing accessible content for the visually impaired.
📚 Dataset Summary
Language: Persian (Farsi)
Speaker: Single female voice
Total Duration: 94.55 hours
Recording Source: Articles from Nasle Mana
Domain: Literary and informational prose
Alignment Checked With:… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/PersianVox_NM.
