CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes5.3k downloads2mo agoHugging Face02Dorsaasgari /vibevoice-quran_persian-single-speakeraudio1K<n<10K1 likes4.9k downloads8d agoHugging Face03Reza2kn /persian-ocr-community-dataset-argilla Persian OCR community dataset - Argilla view Lightweight two-column view for Argilla. image is an HF-hosted asset URL and label is JSON containing spatial OCR objects. image1K<n<10K0 likes4.4k downloads2mo agoHugging Face04hezarai /persian-license-plate-v1 Dataset is downloaded from here which was provided at Amirkabir University of Technology. The dataset is labeled by the authors. Experimental results show that the fine-tuned model works well in Persian License Plate. Usage You can download the dataset easily using HF datasets package in Python: !pip install datasets from datasets import load_dataset dataset = load_dataset("hezarai/persian-license-plate-v1", split="train") # Other splits: validation, test print(dataset[0]) imageimage-to-text1K<n<10K11 likes3.9k downloads2y agoHugging Face05Dorsaasgari /vibevoice-gptinformal_persian-single-speakeraudio1K<n<10K0 likes3.2k downloads8d agoHugging Face06Reza2kn /persian-handwriting-pages-3.69m Persian Handwriting Pages 3.69M 3,690,000 deterministic, densely composed Persian handwriting pages. This expansion uses new random seeds and is complementary to Reza2kn/persian-handwriting-pages-369k, not a repetition of its rendered pages. The public viewer intentionally exposes exactly two columns: image and label. Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification. Source handwriting Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.imageimage-to-text1M<n<10M4 likes2.5k downloads2mo agoHugging Face07Reza2kn /persian-ocr-community-datasetimage10K<n<100K1 likes2.4k downloads2mo agoHugging Face08Thomcles /Persian-Farsi-Speechgated Persian (Farsi) TTS Dataset 🗂️ Dataset Description This dataset is a Persian (Farsi) text-to-speech (TTS) corpus built by concatenating and denoising multiple existing Farsi datasets.It is intended for training and evaluation of speech synthesis (TTS) models in Persian. Since the basic datasets were contaminated with unintelligible audio, I used dnsmos to keep only clean audio (mos_ovr >= 3.0, same value as for the Emilia dataset). The dataset contains two main… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/Persian-Farsi-Speech.audiotext-to-speech100K<n<1M23 likes1.9k downloads29d agoHugging Face09Reza2kn /persian-printed-ocr-3.5m Persian Printed OCR 3.5M A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733 rejected rows are excluded. The viewer exposes exactly image and label. Sources AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0) hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.imageimage-to-text1M<n<10M2 likes1.6k downloads2mo agoHugging Face10Mehdinmz /persian-handwritten-digits Persian Handwritten Digits (Farsi) 80,000 grayscale images of handwritten Persian (Farsi) digits — ۰۱۲۳۴۵۶۷۸۹ — organized as an ImageFolder dataset with 10 classes (0–9), 8,000 images per class. Each image is a 28×28 grayscale PNG of a single digit. Classes Class Count 0 (۰) 8,000 1 (۱) 8,000 2 (۲) 8,000 3 (۳) 8,000 4 (۴) 8,000 5 (۵) 8,000 6 (۶) 8,000 7 (۷) 8,000 8 (۸) 8,000 9 (۹) 8,000 Usage from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Mehdinmz/persian-handwritten-digits.image10K<n<100K1 likes1.2k downloads1mo agoHugging Face11Reza2kn /persian-handwriting-pages-369k Persian Handwriting Pages 369K Full-page Persian handwriting compositions on scanned paper backgrounds. Each row deliberately has only two fields: image: the composed full-page image label: its complete line-separated Persian transcription, ordered from top to bottom The pages are composed from labeled real handwriting crops with page-level ink normalization, controlled RTL layout variation, collision prevention, and exact transcription provenance. The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.imageimage-to-text100K<n<1M4 likes1.2k downloads2mo agoHugging Face12saeedzou /persianvox_2_rawgatedaudio0 likes1.1k downloads27d agoHugging Face13MatinaAI /persian_bbh Persian BBH This is BIG-bench Hard dataset translated to Persian using GPT-4o-mini. We use 19 out of 23 original tasks in BIG-bench Hard. textquestion-answering1K<n<10K5 likes935 downloads2y agoHugging Face14ParsBench /PersianSyntheticQA Persian Synthetic QA Dataset Persian Synthetic QA is a dataset containing 100,000 synthetic questions and answers in Persian, generated using GPT-4o. The dataset is structured as conversations between a user and an assistant, with 2,000 records for each of the 50 different topics. Each conversation consists of messages with two distinct roles: "user" messages containing questions in Persian, and "assistant" messages containing the corresponding answers. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ParsBench/PersianSyntheticQA.textquestion-answering10K<n<100K6 likes796 downloads2y agoHugging Face15sepidmnorozy /Persian_sentimenttext10K<n<100K3 likes733 downloads4y agoHugging Face16saeedzou /persianvox_2_audiogatedaudio1K<n<10K0 likes725 downloads26d agoHugging Face17PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes661 downloads1y agoHugging Face18Arshia82sbn /Finglish-To-Persian-Dataset-Large Finglish to Persian Large Dataset A massive-scale parallel corpus containing over 9.8 million sentence pairs for Finglish (Latin-script Persian) to Persian script transliteration. This dataset provides a robust foundation for training and fine-tuning seq2seq models, normalizing user-generated text, and enhancing Persian input methods. What is Finglish? Finglish (also known as Pinglish) is the practice of writing Persian using the Latin alphabet. Because there is… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Finglish-To-Persian-Dataset-Large.texttranslation10M<n<100M1 likes621 downloads2mo agoHugging Face19Reza2kn /visualears-persian-asr-16k 🗂️ visualears-persian-asr-16k English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Main public Persian ASR audio/text dataset: 3.93M 16 kHz rows. مجموعه‌دادهٔ اصلی و عمومی شنوا برای آموزش بازشناسی گفتار فارسی؛ شامل صوت ۱۶ کیلوهرتز، متن و فرادادهٔ منشأ در مقیاس چندمیلیونی. 🧩 Role flagship training corpus پیکرهٔ اصلی آموزش 📦 Snapshot 176 files; approximately 512.17 GB 176… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-persian-asr-16k.audio1M<n<10M1 likes555 downloads2mo agoHugging Face20Reza2kn /persian-ocr-bench-submitted10-bbox-crops Persian OCR benchmark — selected submitted bbox crops This dataset contains the non-empty OCR bboxes from the ten explicitly selected submitted pages in persian_ocr_bench_bbox_review. Each row is one PNG crop. gold_text is the current editable OCR content from the live Argilla bbox field (content_text). Geometry is stored both as source page pixels and as percentages of the source page. The original record ID, external ID, bbox ID, source URL, and SHA-256 hashes are included for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-bbox-crops.imageimage-to-textn<1K1 likes549 downloads24d agoHugging Face21AlirezaFzp /persian-abusive-words Persian Abusive Words Dataset This is a labeled dataset of Persian Abusive Words, originally sourced from Persian Abusive Words GitHub repository. The dataset has been split by the contributor into two subsets: train and test. This dataset can be used for developing systems to detect and filter offensive or abusive language in various contexts. It is particularly useful for identifying inappropriate words and managing content moderation in applications where Persian language… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaFzp/persian-abusive-words.text10K<n<100K2 likes495 downloads2y agoHugging Face22pymmdrza /Common-Voice-Speech-26.0-Persian-Clean Persian Common Voice Clean Dataset This dataset is a cleaned and prepared subset of the Persian (فارسی - fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12. The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences and come from 3,791 speakers. The original Persian Common Voice release contains… See the full description on the dataset page: https://huggingface.co/datasets/pymmdrza/Common-Voice-Speech-26.0-Persian-Clean.audiotext-to-speech10K<n<100K2 likes482 downloads1mo agoHugging Face23PerSets /filimo-persian-asrThis dataset consists of about 400 hours of audio extracted from various Filimo videos in the Persian language. Note: This dataset contains raw, unvalidated transcriptions. Users are advised to: 1. Perform their own quality assessment 2. Create their own train/validation/test splits based on their specific needs 3. Validate a subset of the data if needed for their use caseautomatic-speech-recognition7 likes437 downloads2y agoHugging Face24SeyedAli /Persian-Text-SentimentDataset Classes negetive :0 positive :1 texttext-classification10K<n<100K4 likes434 downloads3y agoHugging Face25shenasa /bookroom-persian-book-covers-and-titlesimage10K<n<100K3 likes404 downloads1y agoHugging Face26PerSets /youtube-persian-asrThis dataset consists of over 385 hours of audio extracted from various YouTube videos in the Persian language. Note: This dataset contains raw, unvalidated transcriptions. Users are advised to: 1. Perform their own quality assessment 2. Create their own train/validation/test splits based on their specific needs 3. Validate a subset of the data if needed for their use caseautomatic-speech-recognition7 likes388 downloads2y agoHugging Face27MR3z4 /persian-accents-benchmark Persian Accents Benchmark Dataset Summary A benchmark for Persian automatic speech recognition (ASR): 279 short utterances of informal Persian (Farsi) dialect speech across 16 regional accents, released as a fixed evaluation set. Total audio duration is approximately 4.4 hours. The primary label is the transcription; each utterance also carries an accent label (usable for accent classification as a secondary task) and an emotion label as auxiliary metadata. This… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-accents-benchmark.audioautomatic-speech-recognitionn<1K2 likes380 downloads1mo agoHugging Face28asparius /Persian-Food-Sentiment Persian Food Sentiment Dataset This data is orinally from https://hooshvare.github.io/docs/datasets/sa. BibTeX Citation If you use this dataset, please cite following paper: @article{ParsBERT, title={ParsBERT: Transformer-based Model for Persian Language Understanding}, author={Mehrdad Farahani, Mohammad Gharachorloo, Marzieh Farahani, Mohammad Manthouri}, journal={ArXiv}, year={2020}, volume={abs/2005.12515} } texttext-classification10K<n<100K1 likes378 downloads2y agoHugging Face29MohammadReza-Halakoo /Persian_Common_Voice_17_0audio100K<n<1M9 likes373 downloads2y agoHugging Face30MohammadGholizadeh /Tabaghe16_dataset_persianaudio100K<n<1M4 likes364 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.