CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Dorsaasgari /vibevoice-quran_persian-single-speakeraudio1K<n<10K1 likes4.9k downloads10d agoHugging Face02Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.7k downloads2mo agoHugging Face03Reza2kn /persian-ocr-community-dataset-argilla Persian OCR community dataset - Argilla view Lightweight two-column view for Argilla. image is an HF-hosted asset URL and label is JSON containing spatial OCR objects. image1K<n<10K0 likes4.4k downloads2mo agoHugging Face04hezarai /persian-license-plate-v1 Dataset is downloaded from here which was provided at Amirkabir University of Technology. The dataset is labeled by the authors. Experimental results show that the fine-tuned model works well in Persian License Plate. Usage You can download the dataset easily using HF datasets package in Python: !pip install datasets from datasets import load_dataset dataset = load_dataset("hezarai/persian-license-plate-v1", split="train") # Other splits: validation, test print(dataset[0]) imageimage-to-text1K<n<10K11 likes3.6k downloads2y agoHugging Face05Dorsaasgari /vibevoice-gptinformal_persian-single-speakeraudio1K<n<10K0 likes3.2k downloads10d agoHugging Face06Reza2kn /persian-ocr-community-datasetimage10K<n<100K1 likes2.4k downloads2mo agoHugging Face07Reza2kn /persian-handwriting-pages-3.69m Persian Handwriting Pages 3.69M 3,690,000 deterministic, densely composed Persian handwriting pages. This expansion uses new random seeds and is complementary to Reza2kn/persian-handwriting-pages-369k, not a repetition of its rendered pages. The public viewer intentionally exposes exactly two columns: image and label. Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification. Source handwriting Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.imageimage-to-text1M<n<10M4 likes2.3k downloads2mo agoHugging Face08Thomcles /Persian-Farsi-Speechgated Persian (Farsi) TTS Dataset 🗂️ Dataset Description This dataset is a Persian (Farsi) text-to-speech (TTS) corpus built by concatenating and denoising multiple existing Farsi datasets.It is intended for training and evaluation of speech synthesis (TTS) models in Persian. Since the basic datasets were contaminated with unintelligible audio, I used dnsmos to keep only clean audio (mos_ovr >= 3.0, same value as for the Emilia dataset). The dataset contains two main… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/Persian-Farsi-Speech.audiotext-to-speech100K<n<1M24 likes1.9k downloads1mo agoHugging Face09Reza2kn /persian-printed-ocr-3.5m Persian Printed OCR 3.5M A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733 rejected rows are excluded. The viewer exposes exactly image and label. Sources AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0) hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.imageimage-to-text1M<n<10M2 likes1.6k downloads2mo agoHugging Face10Reza2kn /persian-handwriting-pages-369k Persian Handwriting Pages 369K Full-page Persian handwriting compositions on scanned paper backgrounds. Each row deliberately has only two fields: image: the composed full-page image label: its complete line-separated Persian transcription, ordered from top to bottom The pages are composed from labeled real handwriting crops with page-level ink normalization, controlled RTL layout variation, collision prevention, and exact transcription provenance. The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.imageimage-to-text100K<n<1M4 likes1.1k downloads2mo agoHugging Face11MatinaAI /persian_bbh Persian BBH This is BIG-bench Hard dataset translated to Persian using GPT-4o-mini. We use 19 out of 23 original tasks in BIG-bench Hard. textquestion-answering1K<n<10K5 likes966 downloads2y agoHugging Face12ParsBench /PersianSyntheticQA Persian Synthetic QA Dataset Persian Synthetic QA is a dataset containing 100,000 synthetic questions and answers in Persian, generated using GPT-4o. The dataset is structured as conversations between a user and an assistant, with 2,000 records for each of the 50 different topics. Each conversation consists of messages with two distinct roles: "user" messages containing questions in Persian, and "assistant" messages containing the corresponding answers. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ParsBench/PersianSyntheticQA.textquestion-answering10K<n<100K6 likes796 downloads2y agoHugging Face13PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes661 downloads1y agoHugging Face14Reza2kn /persian-ocr-bench-submitted10-bbox-crops Persian OCR benchmark — selected submitted bbox crops This dataset contains the non-empty OCR bboxes from the ten explicitly selected submitted pages in persian_ocr_bench_bbox_review. Each row is one PNG crop. gold_text is the current editable OCR content from the live Argilla bbox field (content_text). Geometry is stored both as source page pixels and as percentages of the source page. The original record ID, external ID, bbox ID, source URL, and SHA-256 hashes are included for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-bbox-crops.imageimage-to-textn<1K1 likes655 downloads26d agoHugging Face15Arshia82sbn /Finglish-To-Persian-Dataset-Large Finglish to Persian Large Dataset A massive-scale parallel corpus containing over 9.8 million sentence pairs for Finglish (Latin-script Persian) to Persian script transliteration. This dataset provides a robust foundation for training and fine-tuning seq2seq models, normalizing user-generated text, and enhancing Persian input methods. What is Finglish? Finglish (also known as Pinglish) is the practice of writing Persian using the Latin alphabet. Because there is… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Finglish-To-Persian-Dataset-Large.texttranslation10M<n<100M1 likes637 downloads2mo agoHugging Face16Reza2kn /visualears-persian-asr-16k 🗂️ visualears-persian-asr-16k English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Main public Persian ASR audio/text dataset: 3.93M 16 kHz rows. مجموعه‌دادهٔ اصلی و عمومی شنوا برای آموزش بازشناسی گفتار فارسی؛ شامل صوت ۱۶ کیلوهرتز، متن و فرادادهٔ منشأ در مقیاس چندمیلیونی. 🧩 Role flagship training corpus پیکرهٔ اصلی آموزش 📦 Snapshot 176 files; approximately 512.17 GB 176… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-persian-asr-16k.audio1M<n<10M1 likes523 downloads2mo agoHugging Face17pymmdrza /Common-Voice-Speech-26.0-Persian-Clean Persian Common Voice Clean Dataset This dataset is a cleaned and prepared subset of the Persian (فارسی - fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12. The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences and come from 3,791 speakers. The original Persian Common Voice release contains… See the full description on the dataset page: https://huggingface.co/datasets/pymmdrza/Common-Voice-Speech-26.0-Persian-Clean.audiotext-to-speech10K<n<100K2 likes478 downloads1mo agoHugging Face18AlirezaFzp /persian-abusive-words Persian Abusive Words Dataset This is a labeled dataset of Persian Abusive Words, originally sourced from Persian Abusive Words GitHub repository. The dataset has been split by the contributor into two subsets: train and test. This dataset can be used for developing systems to detect and filter offensive or abusive language in various contexts. It is particularly useful for identifying inappropriate words and managing content moderation in applications where Persian language… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaFzp/persian-abusive-words.text10K<n<100K2 likes418 downloads2y agoHugging Face19shenasa /bookroom-persian-book-covers-and-titlesimage10K<n<100K3 likes402 downloads1y agoHugging Face20MohammadReza-Halakoo /Persian_Common_Voice_17_0audio100K<n<1M9 likes372 downloads2y agoHugging Face21MR3z4 /persian-accents-benchmark Persian Accents Benchmark Dataset Summary A benchmark for Persian automatic speech recognition (ASR): 279 short utterances of informal Persian (Farsi) dialect speech across 16 regional accents, released as a fixed evaluation set. Total audio duration is approximately 4.4 hours. The primary label is the transcription; each utterance also carries an accent label (usable for accent classification as a secondary task) and an emotion label as auxiliary metadata. This… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-accents-benchmark.audioautomatic-speech-recognitionn<1K2 likes364 downloads1mo agoHugging Face22PersianML /persian-text-corpus Persian Corpus (Merged) Dataset Summary Persian Corpus (Merged) is a large-scale, Persian corpus meticulously aggregated from multiple high-quality Persian datasets available on the Hugging Face Hub. Designed to advance Persian NLP research and applications, this corpus consolidates diverse textual sources into a single resource, providing researchers and developers with a robust foundation for training and evaluating language models. Why Use This… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-text-corpus.texttext-generation10M<n<100M0 likes359 downloads2mo agoHugging Face23MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes356 downloads1y agoHugging Face24MohammadGholizadeh /Tabaghe16_dataset_persianaudio100K<n<1M4 likes352 downloads9mo agoHugging Face25cnababaie /persian-poetry-metersPersian poems with their corresponding meters, from ganjoor.net. texttext-classification100K<n<1M2 likes337 downloads2y agoHugging Face26mohajesmaeili /Persian_Arabic_TextLine_Image_Ocr_Mediumimage100K<n<1M18 likes337 downloads1y agoHugging Face27MahtaFetrat /HomoRich-G2P-Persian HomoRich: A Persian Homograph Dataset for G2P Conversion Overview HomoRich is the first large-scale, sentence-level Persian homograph dataset designed for grapheme-to-phoneme (G2P) conversion tasks. It addresses the scarcity of balanced, contextually annotated homograph data for low-resource languages. The dataset was created using a semi-automated pipeline combining human expertise and LLM-generated samples, as described in the paper:"Fast, Not Fancy: Rethinking G2P… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/HomoRich-G2P-Persian.texttranslation100K<n<1M10 likes328 downloads1y agoHugging Face28SajjadAyoubi /persian_qa\\\\\\\Persian Question Answering (PersianQA) Dataset is a reading comprehension dataset on Persian Wikipedia. The crowd-sourced dataset consists of more than 9,000 entries. Each entry can be either an impossible to answer or a question with one or more answers spanning in the passage (the context) from which the questioner proposed the question. Much like the SQuAD2.0 dataset, the impossible or unanswerable questions can be utilized to create a system which "knows that it doesn't know the answer".text1K<n<10K9 likes317 downloads5y agoHugging Face29kakooch /persian-poetry-qa Persian Poetry Dataset Dataset Description Overview This dataset contains a collection of Persian poems structured in a question-answering format. The dataset is derived from various Persian poets and their poems, providing a rich source for exploring Persian poetry in a structured manner suitable for machine learning applications, especially in natural language processing tasks like question answering. Data Collection Data Collection Source: The… See the full description on the dataset page: https://huggingface.co/datasets/kakooch/persian-poetry-qa.text1M<n<10M7 likes312 downloads3y agoHugging Face30Reza2kn /persian-ocr-gemini37-wins-bina-misses-viewer Gemini 3.7 exact / Bina miss OCR crops 53 non-empty bbox crops from the PersianVLM submitted-10 benchmark where google/gemini-3.7-flash was normalized-exact and Bina Koochik 0.1 was not. This is a minimal Hugging Face ImageFolder dataset for reliable viewer support. It contains exactly two columns: image and ocr. The ocr value is Gemini's actual output for the corresponding crop. imagen<1K0 likes297 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.