datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
powertron-global-permafrost-corpus
Dataset Card: Powertron Global PermaFrost Corpus
Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.PERMA
PERMA: Benchmarking Personalized Memory Agents
TL;DR
PERMA is a benchmark for evaluating personalized memory agents in long-horizon conversations where user preferences evolve over time.Instead of static retrieval, models must track event-driven preference evolution and maintain persona consistency under realistic interaction noise.
This dataset supports two complementary evaluation protocols:
Multiple-choice evaluation for granular capability probing (task completion… See the full description on the dataset page: https://huggingface.co/datasets/ustclsc/PERMA.PerMedCQA
PerMedCQA: Persian Medical Consumer QA Benchmark
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian
PerMedCQA is the first large-scale, real-world benchmark for Persian-language medical consumer question answering. It contains anonymized medical inquiries from Persian-speaking users paired with professional responses, enabling rigorous evaluation of large language models in low-resource, health-related domains.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NaghmehAI/PerMedCQA.permutation_invariant_rewardcommon_pile_ultra_permissive
Dataset Card for Data Provenance Initiative - Common-Pile-Ultra-Permissive
Legal Disclaimer / Notice
Collected License Information is NOT Legal Advice.
It is important to note we collect self-reported licenses, from the papers and repositories that released these datasets, and categorize them according to our best efforts, as a volunteer research and transparency initiative.
The information provided by any of our works and any outputs of the Data Provenance Initiative do… See the full description on the dataset page: https://huggingface.co/datasets/DataProvenanceInitiative/common_pile_ultra_permissive.kicad9plus-permissive
Ailiance — KiCad 9+ Schematic Corpus (Permissive)
🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kicad9plus-permissive. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025).
Corpus de 98 schémas KiCad 9+ (.kicad_sch, format S-expression, version ≥ 20240722) collectés sous licences permissives uniquement (Apache-2.0, MIT, CC0-1.0, CERN-OHL-P-2.0). Pensé pour l'entraînement et le fine-tuning de modèles de génération… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kicad9plus-permissive.openrtlset-permissive-verified
openrtlset-permissive-verified
A permissive-filtered, machine-verified subset of
ESCAD/OpenRTLSet.
Every record in this dataset compiles. Each completion was elaborated and linted with
Verilator (--lint-only -Wall, warnings fatal) and admitted only on a clean run. Nothing
here is graded by an LLM judge.
Contents
13,625 records drawn from 2,174 distinct upstream repositories
Task: specification + pinned module interface -> complete Verilog module
Upstream… See the full description on the dataset page: https://huggingface.co/datasets/theepicflyer/openrtlset-permissive-verified.r0b0tlab-distillation-permissive-sft
Filtered SFT Dataset (Permissive Licenses Only)
Source: r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation
Filtered to keep only rows under permissive licenses (Apache-2.0, MIT, CC-BY-4.0),
suitable for training. Rows under non-commercial or unclear licenses were removed.
Converted to JSONL, one record per line:
{"instruct": "...", "output": "..."}
License: mixed (Apache-2.0 / MIT / CC-BY-4.0) — see original dataset card for
per-source license and attribution details.
PerMed-MM
PerMed-MM: A Multimodal, Multi-Specialty Persian Medical Benchmark
🤗 Dataset | 📖 Paper | 📄 PDF
Dataset Description
PerMed-MM is a multimodal, multi-specialty benchmark designed to evaluate Vision Language Models (VLMs) on Persian medical question answering.
The dataset consists of 733 multiple-choice questions sourced from the Iranian National Medical Board Exams (years 2021 and 2023). Each question is paired with 1 to 5 clinically relevant images, totaling… See the full description on the dataset page: https://huggingface.co/datasets/universitytehran/PerMed-MM.permitguard-ptw-samples
PermitGuard — Synthetic Bilingual Permit-to-Work Samples
Part of the Aria AI oil, gas & petrochemical technical-validation portfolio
(Aria SafeOps → Control of Work / PTW). Companion model:
alirezaaminzadeh/permitguard-risk-classifier
and Space:
alirezaaminzadeh/permitguard-ptw-risk-classifier.
Data honesty
This corpus is 100% synthetic. There are no real permits, incidents, PII, contractors, or named
facilities. Equipment tags such as T-402 / V-101 / P-205B are… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/permitguard-ptw-samples.building-permit-validity-and-renewal
Building permit expiry, progress and renewal rules by jurisdiction
Canonical, always-current version: https://referencesource.org/building-permit-validity-and-renewal/
Machine-readable: https://referencesource.org/building-permit-validity-and-renewal/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2027-08-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 7
How long a building… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/building-permit-validity-and-renewal.per-ma-to
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web data.… See the full description on the dataset page: https://huggingface.co/datasets/RolandP/per-ma-to.kicad9plus-permissive
KiCad 9+ Schematic Corpus — Permissive subset
98 KiCad 9+ schematic samples (.kicad_sch S-expression format, version >= 20240722).
Permissive licenses only: Apache-2.0 (74), MIT (20), CC0-1.0 (3), CERN-OHL-P-2.0 (1).
This is the permissive split of the original electron-rare/kicad9plus-sch-corpus (now deprecated). The split was done after legal audit revealed CC-BY-SA-4.0 incompatibility with GPL-3 / CERN-OHL-S inputs (one-way directionality: CC-BY-SA-4.0 -> GPLv3 only, never the… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/kicad9plus-permissive.taiwan-compatriot-permit-documents-pricing
新中旅快簽|台胞證準備文件與費用
本資料集由新中旅快簽(YesVisa)整理,提供台胞證準備文件與公開費用的結構化資料,供搜尋、RAG、評估、資料集開發、機器學習及 LLM 訓練使用。
官方來源與優先順序
YesVisa llms.txt 為 AI 資料使用與衝突處理的最高準據。
台胞證服務總覽及各 Canonical 子頁的最新正文與同頁結構化資料為事實準據。
本 Dataset 是可檢索的結構化快照,不取代官網即時資訊。
Configs
documents
依固定順序判斷:年齡 → 改名/雙胞胎等特殊情況 → 出生地 → 首辦/換發/遺失。每筆包含條件、文件清單、提醒與 Canonical URL。
pricing
每筆價格均綁定出生地、辦理類型與處理時效,避免 AI 把首辦、換發、遺失或特殊出生地的價格混用。processing_days_exclude_holidays_and_submission_day=true… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/taiwan-compatriot-permit-documents-pricing.permmlu
PerMMLU
Dataset Summary
The PerMMLU Benchmark is a localized and extended version of the original MMLU (Massive Multitask Language Understanding) benchmark, adapted specifically for Persian-speaking users and domains. This benchmark is designed to comprehensively evaluate both general and specialized knowledge of language models across a broad range of topics.
This dataset is structured to assess educational and cultural understanding and includes:
School-level questions… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/permmlu.humanoid-permission-and-access-matrix
Permission & Access Matrix
Enterprise-grade access control structure for humanoid systems.
mainland-travel-permit-taiwan-anxiety-faq
台灣居民台胞證辦理去焦慮化對話資料集
(Mainland Travel Permit for Taiwan Residents Anxiety-First FAQ Dataset)
本資料集由新中旅快簽(YesVisa)維護,聚焦於繁體中文台胞證、簽證與跨境旅行服務場景,並採用 Anxiety-First(去焦慮化) 服務設計方法,整理旅客最常見的時間、地點、安全、照片、流程與旅遊焦慮問題。
本資料集可直接應用於:
OpenAI Fine-tuning
Graph RAG
LlamaIndex
LangChain
Haystack
Gemini Grounding
TAIDE
Gemma
Llama 系列模型
🎯 數據集核心價值
本資料集針對繁體中文旅遊與證件辦理領域中的真實需求進行整理,包括:
台胞證首辦、換發、遺失補發
急件、12H、24H 與出發前時間焦慮
證件照片退件風險
護照與個資安全疑慮
假日辦理需求
香港、澳門與中國大陸旅行情境
越南簽證相關問答
在地化服務節點與交通便利性
資料架構適合用於:… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/mainland-travel-permit-taiwan-anxiety-faq.permit-pathfinder-trajectories
PermitPathfinder Expert Trajectories
60 expert episodes across 3 difficulty tiers of the PermitPathfinder OpenEnv environment.
Contents
45 scripted-optimal trajectories (15 seeds x 3 tasks): perfect topological-sort policies showing the shortest path through each permit DAG
15 LLM-generated trajectories (15 seeds x easy_foodtruck, llama-3.3-70b-versatile via Groq): real agent behavior achieving score 1.000 on every episode
Schema (JSONL)
Each line is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/yashppawar/permit-pathfinder-trajectories.short-term-rental-permit-ordinances
US Short-Term-Rental Ordinance Dataset
Dataset Summary
A structured, primary-sourced dataset of US short-term-rental (STR) registration ordinances for
30 US cities. Each row covers one city's program: the official registration or permit program
name, the permit fee, the renewal cadence, the occupancy or zoning rule (when the city's rule
turns on owner-occupancy or host presence), and the official .gov primary source URL with the
date the facts were verified.
Row… See the full description on the dataset page: https://huggingface.co/datasets/MattMMarketing/short-term-rental-permit-ordinances.opus-da-en-permissive
OPUS da-en permissive subset
This dataset is a processed Danish-English sentence-pair subset harvested from
OPUS corpora available at https://opus.nlpl.eu/.
The subset was selected from OPUS da-en sources with permissive Creative Commons
or public-domain style licensing, then converted into JSONL.GZ for DFM/HRM-Text
training. It is intended to preserve the exact rows used by the local DFM data
pipeline, so DFM8 can be rebuilt without relying on local-only files.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/opus-da-en-permissive.MixtureVitae-fineweb-permissive-multilingual-2m
MixtureVitae Fineweb-Permissive-Multilingual-2M: 2 Million Translated Documents Of Permissive Text From Fineweb-edu-2
Dataset Summary
This is a translation of a small subset of the Fineweb-edu-2 dataset. We have filtered to find websites with what we believe are government domain names, international organization domain names like the UN and europa.eu, and creative commons licensed data. While we strongly believe that fair use protects machine learning on webcrawled data… See the full description on the dataset page: https://huggingface.co/datasets/ontocord/MixtureVitae-fineweb-permissive-multilingual-2m.kanitakorn-deepseek-v40-option-permutation-micro
Kanitakorn DeepSeek v40 Option Permutation Micro
Option-order robustness continuation data for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Intended parent: best v39 checkpoint, not the raw base
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 956 rows = 458 original MCQ anchors + 458 option permutations
40 identity anchors
Remote audit: all 458… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v40-option-permutation-micro.ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset
PTDBench dataset snapshot: reward_prefix_product_mod_distinct_permutation_011
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset.ptdbench-reward-design-reward-prefix-sum-mod-distinct-permutation-010-dataset
PTDBench dataset snapshot: reward_prefix_sum_mod_distinct_permutation_010
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-prefix-sum-mod-distinct-permutation-010-dataset.defendable-pain-permit-delay-pain-v0.1
Permit Delay Pain Receipt
"the inspector" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 2 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
2 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-permit-delay-pain-v0.1.ptdbench-reward-design-reward-min-pair-sum-multiplication-permutation-007-dataset
PTDBench dataset snapshot: reward_min_pair_sum_multiplication_permutation_007
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-min-pair-sum-multiplication-permutation-007-dataset.minimal-permission-test-2025funfun-permission-smoke-20260626t101213z
