CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MatinaAI /persian_bbh Persian BBH This is BIG-bench Hard dataset translated to Persian using GPT-4o-mini. We use 19 out of 23 original tasks in BIG-bench Hard. textquestion-answering1K<n<10K5 likes1k downloads2y agoHugging Face02ParsBench /PersianSyntheticQA Persian Synthetic QA Dataset Persian Synthetic QA is a dataset containing 100,000 synthetic questions and answers in Persian, generated using GPT-4o. The dataset is structured as conversations between a user and an assistant, with 2,000 records for each of the 50 different topics. Each conversation consists of messages with two distinct roles: "user" messages containing questions in Persian, and "assistant" messages containing the corresponding answers. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/ParsBench/PersianSyntheticQA.textquestion-answering10K<n<100K6 likes794 downloads2y agoHugging Face03PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes652 downloads1y agoHugging Face04MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes639 downloads1y agoHugging Face05cnababaie /persian-poetry-metersPersian poems with their corresponding meters, from ganjoor.net. texttext-classification100K<n<1M2 likes380 downloads2y agoHugging Face06rmoham05 /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.textquestion-answering100K<n<1M0 likes235 downloads2mo agoHugging Face07mainkilora /PersianSyntheticQA Persian Synthetic QA Dataset Persian Synthetic QA is a dataset containing 100,000 synthetic questions and answers in Persian, generated using GPT-4o. The dataset is structured as conversations between a user and an assistant, with 2,000 records for each of the 50 different topics. Each conversation consists of messages with two distinct roles: "user" messages containing questions in Persian, and "assistant" messages containing the corresponding answers. The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/PersianSyntheticQA.textquestion-answering10K<n<100K0 likes216 downloads8mo agoHugging Face08saeid1999 /persian-poetics-kb Persian Poetics Knowledge Base — پایگاه دانش شعر و رپ فارسی مجموعهٔ بازیابی برای ساخت و ارزیابی شعر/رپ فارسی به‌مثابهٔ پرسش‌وپاسخِ محدودیت‌دار و مبتنی بر شواهد (PersianPoet-RAG v2). هر واحد هم حاشیه‌نویسی معنایی (معنا، دامنه، تصویر) دارد و هم آوایی (واج‌ها، ساخت هجا، تکیه، کلید قافیهٔ سخت‌گیرانه/آسان‌گیر، زنجیرهٔ واکه‌ها) — چیزی که بازیاب را قادر می‌کند وزن و قافیه را قبل از فراخوانی مدل زبانی تأمین کند. چه چیزهایی داخل این دیتاست است؟ (What's inside) کانفیگ… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/persian-poetics-kb.tabularquestion-answering10K<n<100K0 likes195 downloads8d agoHugging Face09xmanii /Maux-Persian-SFT-30k Maux-Persian-SFT-30k Dataset Description This dataset contains 30,000 high-quality Persian (Farsi) conversations for supervised fine-tuning (SFT) of conversational AI models. The dataset combines multiple sources to provide diverse, natural Persian conversations covering various topics and interaction patterns. Dataset Structure Each entry contains: messages: List of conversation messages with role (user/assistant/system) and content source: Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/Maux-Persian-SFT-30k.tabularquestion-answering10K<n<100K4 likes129 downloads1y agoHugging Face10PerSets /clinical-persian-qa-ii Clinical Question Answering Dataset II (Farsi) This dataset contains more than 211k questions and more than 700k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset is NOT a part of Clinical Question Answering I dataset and is a whole… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-ii.textquestion-answering100K<n<1M3 likes76 downloads1y agoHugging Face11hamidsalimi /Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1 Persian Civil Procedure QA Dataset Dataset Description این مجموعه‌داده شامل پرسش‌وپاسخ‌های حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است. هر نمونه شامل سه فیلد اصلی است: question: پرسش حقوقی answer: پاسخ پرسش evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است هدف مجموعه‌داده، فراهم‌کردن داده‌ای ساختاریافته برای آموزش، ارزیابی و توسعه مدل‌های زبانی فارسی در زمینه پرسش‌وپاسخ حقوقی است. Dataset Structure نمونه‌ای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.textquestion-answeringn<1K1 likes73 downloads2mo agoHugging Face12Jamalianpour /persian-wikipedia-instruct Persian Wikipedia Instruct A Persian-language instruction-tuning dataset of ~120,000 samples, generated from the Persian Wikipedia (fawiki) article dump. Each article was cleaned to Markdown, chunked by section, and passed to a locally-run Gemma4 model that produced grounded instruction/response pairs across five task types. The result is ready for supervised fine-tuning (SFT) of Persian LLMs. Heads-up: this is synthetic data. The instructions and answers were written by an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Jamalianpour/persian-wikipedia-instruct.texttext-generation100K<n<1M0 likes71 downloads3mo agoHugging Face13MohammadJRanjbar /PersianMedQAgated PersianMedQA PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark A large-scale, expert-validated multiple-choice question set covering 23 medical specialties, collected from 14 years (2011–2024) of Iranian national residency and pre-residency board examinations administered by Sanjeshp (the Medical Education Assessment Center, under the Iranian Ministry of Health). Total items: 20,785 Train 14,549 · Validation 1,000… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA.tabularquestion-answering10K<n<100K8 likes69 downloads2mo agoHugging Face14mshojaei77 /persian-gk-cleanedThis is a cleaned and validated version of the original mshojaei77/persian-gk dataset. The purpose of this version is to ensure robust compatibility with modern fine-tuning workflows that rely on strict chat templates (e.g., tokenizer.apply_chat_template). The cleaning process resolves structural errors in the original dataset that could cause TemplateError or other silent failures during training with models like Gemma 3N, Llama 3, and others. Cleaning and Validation Process… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-gk-cleaned.textquestion-answering1K<n<10K1 likes69 downloads1y agoHugging Face15PerSets /crossword-puzzle-persian-cheat Crossword Puzzle Cheat Dataset (Persian) This dataset consists of 30157 pairs of questions and answers. Dataset Description The reference for this dataset is jadvalyab.ir website. Usage Huggingface datasets library: from datasets import load_dataset dataset = load_dataset('PerSets/crossword-puzzle-persian-cheat') License CC0-v1.0 textquestion-answering10K<n<100K1 likes68 downloads2y agoHugging Face16MCINext /persian-nlu Persian NLU Dataset Summary The Persian NLU Benchmark is a curated collection of existing Persian datasets designed to evaluate Natural Language Understanding (NLU) capabilities across a diverse range of tasks. It provides a unified benchmark suite to assess different cognitive aspects of large language models (LLMs) in Persian. This benchmark includes the following tasks and datasets: Text Classification: Synthetic Persian Tone SID Natural Language Inference… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-nlu.text-classification0 likes67 downloads1y agoHugging Face17PerSets /clinical-persian-qa-i Clinical Question Answering Dataset I (Farsi) This dataset contains approximately 50k questions and around 60k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset is NOT a part of Clinical Question Answering II dataset and is a complete… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-i.textquestion-answering10K<n<100K2 likes62 downloads1y agoHugging Face18AliMoameri /drhast-persian-medical-QA Medical Question and Answer Dataset Crawled from https://drhast.com. Dataset Structure { "title": "پایین امدن دیابت بارداری و ادامه دادن قرص گلوکوفاژ؟", "url": "https://drhast.com/q/yLvD6", "text": "سلام من هفته ۳۳ بارداری هستم هفته ۳۱ ازمایش دادم تشخیص دیابت دادن ناشتا۹۳ یکساعت بعد خوردن محلول ۲۰۷ و دوساعت بعد۱۷۲ متخصص داخلی روزی دوتا گلوکوفاژ و یه سری محدودیتا تغذیه مشخص کردن دوکیلو وزن کم کردم و امروز مجدد ازمایش دادم قند ناشتا ۸۴ و دوساعت بعد… See the full description on the dataset page: https://huggingface.co/datasets/AliMoameri/drhast-persian-medical-QA.textquestion-answering1K<n<10K2 likes57 downloads2y agoHugging Face19aictsharif /persian-med-qa 🏥 Persian Medical Question Answering Dataset Dataset Summary The Persian Medical QA Dataset is a high-quality, expert-curated collection of question–answer (QA) pairs in Persian (Farsi), designed for developing and evaluating natural language processing (NLP) models for medical question answering. All answers are grounded in reliable medical resources, including standard medical textbooks (e.g., Harrison’s Principles of Internal Medicine) and authoritative medical… See the full description on the dataset page: https://huggingface.co/datasets/aictsharif/persian-med-qa.textquestion-answering100K<n<1M2 likes53 downloads1y agoHugging Face20yoyo-research-group /Persian_Riddles Persian Riddles (چیستان‌های فارسی) A curated collection of 1,065 Persian-language riddles, each paired with its answer. The set spans classic Persian riddles (چیستان), wordplay and letter puzzles, lateral-thinking brain-teasers, and knowledge-based trick questions gathered from a range of Persian sources. Duplicates were removed during compilation: entries were normalized for Persian/Arabic character variants (ی/ي, ک/ك), zero-width joiners, diacritics, and punctuation, and a… See the full description on the dataset page: https://huggingface.co/datasets/yoyo-research-group/Persian_Riddles.textquestion-answering1K<n<10K1 likes49 downloads2mo agoHugging Face21sosa123454321 /persian-legal-kb Persian Legal Knowledge Base — پایگاه دانش حقوقی فارسی A provenance-first Persian legal corpus built for retrieval-augmented and knowledge-graph (KAG) systems, not for fine-tuning knowledge into weights. Every record answers four questions a legal AI must never get wrong: is this law or just a bill? · is it still in force? · which article? · who said so? Why this dataset exists Iranian law moved in the two weeks before this dataset was built: قانون ملی توسعه هوش… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/persian-legal-kb.question-answeringn<1K0 likes47 downloads11d agoHugging Face22persiannlp /parsinlu_reading_comprehensionA Persian reading comprehenion task (generating an answer, given a question and a context paragraph). The questions are mined using Google auto-complete, their answers and the corresponding evidence documents are manually annotated by native speakers.question-answering1K<n<10K1 likes44 downloads4y agoHugging Face23mshojaei77 /persian-gk Dataset Card for persian-gk (Persian General Knowledge) Dataset Summary persian-gk is a cleaned and structured collection of Persian (Farsi) conversation pairs covering a wide range of general-knowledge topics. Each conversation is formatted in ChatML style with explicit system, user, and assistant roles, enabling straightforward use for both instruction-tuning and chat-style language-model training. Language: Persian (fa) Size: 5 897 conversations, 2–8 turns each (≈… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-gk.textquestion-answering1K<n<10K1 likes44 downloads1y agoHugging Face24ZharfaTech /ZharfaTech-Open-Platypus-Persian-Farsi Persian Open-Platypus About ZharfaTech ZharfaTech is a pioneer in developing Language Learning Models (LLMs) tailored for the Persian language, aiming to empower over 100 million Persian speakers worldwide. Our mission encompasses bridging the digital divide in LLM-related services like content generation, customer relationship systems, and more, with a dual approach of fostering open-source collaboration and delivering high-value, specialized closed-source solutions.… See the full description on the dataset page: https://huggingface.co/datasets/ZharfaTech/ZharfaTech-Open-Platypus-Persian-Farsi.texttext-generation10K<n<100K2 likes42 downloads3y agoHugging Face25arshiaafshani /persian-natural-fluently Persian scientific dataset I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology. The content of the dataset has been generated by : Human, Grok3, DeepSeek R1. License This dataset is licensed under apache-2.0. texttext-generationn<1K14 likes41 downloads1y agoHugging Face26mainkilora /mauxi-COT-Persian 🧠 mauxi-COT-Persian Dataset Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform 🌟 Overview mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/mauxi-COT-Persian.textquestion-answeringn<1K0 likes38 downloads8mo agoHugging Face27artindnr /Persian-Thinking Persian-Thinking Persian-Thinking is a small Persian-language reasoning/thinking dataset created by sampling and translating a subset of SmolTalk2. Dataset Details 1,000 samples (999 after processing) drawn from the smoltalk_systemchats_Qwen3_32B_think subset of SmolTalk2, part of its SFT split. That source subset consists of system-chat conversations generated with Qwen3-32B in thinking mode, meaning each assistant response includes an explicit reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/Persian-Thinking.texttext-generationn<1K3 likes36 downloads2mo agoHugging Face28Mobinatj /PersianMHQAThe PersianMHQA Dataset is the first open-domain, multi-hop question answering dataset, containing 7,000 questions generated using Persian Wikipedia texts. We have currently made available 500 samples from the dataset records to introduce this dataset. Soon, after the publication of the paper, the entire dataset will be accessible. Contact Us ** Mobina Taji Email 📧**: mtaji3405@gmail.com question-answering0 likes34 downloads2y agoHugging Face29ErfanMoosaviMonazzah /ErfanClone-persian-alpaca Persian Alpaca This dataset is directly cloned from this one. All the credits should go to iamshnoo for translating alpaca to persian. Translated from yahma/alpaca-cleaned using NLLB-1.3B texttext-generation10K<n<100K1 likes33 downloads2y agoHugging Face30ParsBench /Persian-MuSR Persian MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning on Persian Language This is the Persian-translated version (using GPT-4o) of the original dataset MuSR. Acknowledgments Special thanks to AvalAI for sponsoring this project through their AvalAward program This dataset was made possible by AvalAI's generous support and commitment to advancing Persian language AI research textquestion-answeringn<1K0 likes33 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.