datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.Iranian_olympiad_of_informatics_multimodal_questionsiran-turkiye-startup-landing-kb
Iran → Türkiye Startup Landing — legal/migration knowledge base
One dataset for the whole platform (rule: one dataset, one space, one vector
index — never several). Nightly snapshots produced by rag/scrape.py:
news, academic, official (göç idaresi / ministry), legislation and directory
sources about Iranian founders landing startups in Türkiye.
All PII is scrubbed at fetch time (rag/fetch.scrub_pii).
Chunks are stored as JSONL per snapshot day under data/kb/<date>/.
This… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/iran-turkiye-startup-landing-kb.IRAN-MADANI-LAWiranian-elderly-psychospiritual-interviews
Iranian Elderly Psycho-Spiritual Interviews
A culturally grounded, fully synthetic conversational interview dataset for assessing the mental and spiritual health of Iranian older adults, generated using Large Language Models.
Dataset Summary
This dataset introduces a culturally grounded, fully synthetic conversational interview corpus designed for the assessment and analysis of mental and spiritual health among Iranian older adults. All interviews are conducted in… See the full description on the dataset page: https://huggingface.co/datasets/liamirali/iranian-elderly-psychospiritual-interviews.Iranian-Court-Rulings-25K
Iranian Court Rulings Dataset
A structured Persian-language corpus containing 25,617 Iranian judicial rulings collected from publicly accessible pages of the Iranian National Judicial Opinions database.
The dataset is intended for research and development in Persian Legal NLP, Information Retrieval, Retrieval-Augmented Generation (RAG), semantic search, legal document understanding, and related areas.
Dataset Overview
Number of records: 25,617
Language: Persian… See the full description on the dataset page: https://huggingface.co/datasets/sinamahallati/Iranian-Court-Rulings-25K.
