datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wdc-common-crawl-embedded-jsonldgithub-code-permissive-sampleSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'].It is intended to be used for training code language classifier.
finephrase_permissivepowertron-global-permafrost-corpus
Dataset Card: Powertron Global PermaFrost Corpus
Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.moral_education_permissive
Moral Education Permissive
Permissive-source subset of locuslab/moral_education, filtered using row-level metadata.url with the local permissiveness rules in this workspace. The output preserves the original row fields and adds url, idx, dump, language, source_config, is_permissive, is_oss, and permissive_reason where applicable.
Source configs included: score_4_morals, score_5_morals. Source split: train.
Rows scanned: 2,806,450. Rows kept: 72,943. Keep rate: 2.5991%.
See… See the full description on the dataset page: https://huggingface.co/datasets/ontocord/moral_education_permissive.c4-bbc-news
Dataset Card for BBC News from C4
This dataset provides a filtered subset of BBC News articles from the realnewslike subset of the C4 dataset, containing approximately 77k articles from BBC News domains.
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/permutans/c4-bbc-news
Source Dataset: allenai/c4 (realnewslike subset)
Paper: https://arxiv.org/abs/1910.10683 (C4 paper)
Uses
Direct Use
Suitable for text… See the full description on the dataset page: https://huggingface.co/datasets/permutans/c4-bbc-news.Word_Permuations_EnglishPERMA
PERMA: Benchmarking Personalized Memory Agents
TL;DR
PERMA is a benchmark for evaluating personalized memory agents in long-horizon conversations where user preferences evolve over time.Instead of static retrieval, models must track event-driven preference evolution and maintain persona consistency under realistic interaction noise.
This dataset supports two complementary evaluation protocols:
Multiple-choice evaluation for granular capability probing (task completion… See the full description on the dataset page: https://huggingface.co/datasets/ustclsc/PERMA.ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB
ExpansionRx-OpenADMET Caco-2 Permeability Papp A>B
Caco-2 Permeability Papp A>B dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict Caco-2 Permeability Papp A>B of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3773
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB.atlas-16-verifier-permission-prompt-ablation
ATLAS report 16: does the orchestrator's "cannot solve" clause suppress candidate verification?
Complete raw products of the ATLAS rl-training report 16 experiment
(GitHub issue #36). Two system-prompt arms of the same model over the
same 78 fixed states, greedy decoding, one shared vLLM server.
What the experiment did
The ATLAS orchestrator's frozen system prompt contains the clause
You cannot solve the problem yourself; you decide when to explore
further and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-16-verifier-permission-prompt-ablation.freesound-commercially-permissive-subset-with-captionsExpansionRx_OpenADMET_Caco-2_Permeability_Efflux
ExpansionRx-OpenADMET Caco-2 Permeability Efflux
Caco-2 Permeability Efflux dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict Caco-2 Permeability Efflux of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3777
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Efflux.dummy-lang-subset-dataset-1m-chunksDummy dataset uploaded to test the process, benefits of and constraints on/drawbacks to uploading subsets by a partition of a dataset.
This dataset is made up of fake data to illustrate a long tail of rarer languages similar to Wikipedia/Wikidata's distribution.
The dataset metadata was written automatically by the datasets library.
The config name is the language, and we iterate over all languages to do this.
Since the data is synthetic, we have the list of languages as a variable, without… See the full description on the dataset page: https://huggingface.co/datasets/permutans/dummy-lang-subset-dataset-1m-chunks.PerMedCQA
PerMedCQA: Persian Medical Consumer QA Benchmark
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian
PerMedCQA is the first large-scale, real-world benchmark for Persian-language medical consumer question answering. It contains anonymized medical inquiries from Persian-speaking users paired with professional responses, enabling rigorous evaluation of large language models in low-resource, health-related domains.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NaghmehAI/PerMedCQA.permutation_invariant_rewardcommon_pile_ultra_permissive
Dataset Card for Data Provenance Initiative - Common-Pile-Ultra-Permissive
Legal Disclaimer / Notice
Collected License Information is NOT Legal Advice.
It is important to note we collect self-reported licenses, from the papers and repositories that released these datasets, and categorize them according to our best efforts, as a volunteer research and transparency initiative.
The information provided by any of our works and any outputs of the Data Provenance Initiative do… See the full description on the dataset page: https://huggingface.co/datasets/DataProvenanceInitiative/common_pile_ultra_permissive.theorem-search-dataset-permissive
Theorem Search Dataset
The largest open corpus of informal mathematical theorems: 1,239,720 theorem statements with natural-language slogans from 197,889 papers, designed for semantic theorem retrieval.
Paper: Semantic Search over 9 Million Mathematical Theorems
Demo: huggingface.co/spaces/uw-math-ai/theorem-search
Benchmark results
On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms all… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset-permissive.kicad9plus-permissive
Ailiance — KiCad 9+ Schematic Corpus (Permissive)
🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kicad9plus-permissive. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025).
Corpus de 98 schémas KiCad 9+ (.kicad_sch, format S-expression, version ≥ 20240722) collectés sous licences permissives uniquement (Apache-2.0, MIT, CC0-1.0, CERN-OHL-P-2.0). Pensé pour l'entraînement et le fine-tuning de modèles de génération… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kicad9plus-permissive.Word_Permuations_Englishpermission-command-corpus
Permission Command Corpus
Three views of a command-safety corpus, for local command-risk classification in
front of an LLM or a tool bridge.
gold: trusted rows, 321 in total across three splits
silver_weak_labels: mined weak-label rows from Sigma, LOLBAS, GTFOBins,
Atomic Red Team and Falco, kept as useful but not promoted to gold
review_queue: unresolved rows that should not be treated as trusted
training data
Read this before training on it
Findings from 30… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/permission-command-corpus.us-building-permits
PermitBase — U.S. Residential Building Permits 1980–2024
The most comprehensive historical residential building permit dataset available at the place level.
Place-level | Annual | 1980–2024 | SF/MF differentiated | 51 jurisdictions | 683,986 records
Dataset Description
This dataset contains annual residential building permit data for permit-issuing places
(cities, towns, and unincorporated county areas) across the United States, covering
1980 through 2024. It is derived… See the full description on the dataset page: https://huggingface.co/datasets/thanna94/us-building-permits.serbian-llm-eval
Serbian LLM Eval
This dataset is a republish of
gordicaleksa/serbian-llm-eval-v1
as plain Parquet configs, so it can be loaded with modern versions of the
datasets library (the upstream repo ships a Python loading script that
datasets can no longer execute). Row contents are otherwise unchanged; only
the example_id field is synthesized where the upstream data has no unique
identifier, and the triviaqa answer struct is flattened into
answer_value / answer_aliases columns.… See the full description on the dataset page: https://huggingface.co/datasets/permitt/serbian-llm-eval.PineScripts-Permissive
PineScripts-Permissive
A dataset of Pine Script™ scripts with premissive licenses from TradingView.
oracle_perm_medm_datapermutation-pools
Permutation Pools
Permutation-augmentation pools: demonstration-order permutations of the label-transfer and evidence-tracking sets, used to enlarge filtered test sets to 1000 rows with balanced golds.
Part of the icl-heads collection: the datasets behind a study of which
Llama-3.1-8B-Instruct attention heads mediate in-context evidence accumulation
and in-context label mapping, discovered with Differentiable Circuit Masking
(a learned per-head mask over K/V activations patched… See the full description on the dataset page: https://huggingface.co/datasets/icl-heads/permutation-pools.openrtlset-permissive-verified
openrtlset-permissive-verified
A permissive-filtered, machine-verified subset of
ESCAD/OpenRTLSet.
Every record in this dataset compiles. Each completion was elaborated and linted with
Verilator (--lint-only -Wall, warnings fatal) and admitted only on a clean run. Nothing
here is graded by an LLM judge.
Contents
13,625 records drawn from 2,174 distinct upstream repositories
Task: specification + pinned module interface -> complete Verilog module
Upstream… See the full description on the dataset page: https://huggingface.co/datasets/theepicflyer/openrtlset-permissive-verified.r0b0tlab-distillation-permissive-sft
Filtered SFT Dataset (Permissive Licenses Only)
Source: r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation
Filtered to keep only rows under permissive licenses (Apache-2.0, MIT, CC-BY-4.0),
suitable for training. Rows under non-commercial or unclear licenses were removed.
Converted to JSONL, one record per line:
{"instruct": "...", "output": "..."}
License: mixed (Apache-2.0 / MIT / CC-BY-4.0) — see original dataset card for
per-source license and attribution details.
assay-transfer-record-level-v23-1-bbb-martins-passive_permeability-intern
BBB V23.1: passive_permeability
Exact passive_permeability row filter of pinned BBB V23 revision d49645551057f0dfc562dd7bf2ef7d77ff975a6f. Prompts, targets, candidate order, and split membership are unchanged. The complete parent calibration artifact is retained. V23 has no test split.
train rows: 46,541
validation ranking rows: 1,525
PerMed-MM
PerMed-MM: A Multimodal, Multi-Specialty Persian Medical Benchmark
🤗 Dataset | 📖 Paper | 📄 PDF
Dataset Description
PerMed-MM is a multimodal, multi-specialty benchmark designed to evaluate Vision Language Models (VLMs) on Persian medical question answering.
The dataset consists of 733 multiple-choice questions sourced from the Iranian National Medical Board Exams (years 2021 and 2023). Each question is paired with 1 to 5 clinically relevant images, totaling… See the full description on the dataset page: https://huggingface.co/datasets/universitytehran/PerMed-MM.superglue-hr
BalkanBench SuperGLUE - Croatian
Part of BalkanBench - the open, reproducible
benchmark for language models across Serbian, Croatian, Montenegrin, and
Bosnian (BCMS). Live leaderboard at https://balkanbench.com/leaderboard.
Background: Release of BalkanBench - the vision behind it
(Medium, 2026-04-27).
This is the Croatian SuperGLUE released preview track of BalkanBench v0.1.
Croatian publishes alongside Serbian (the official frozen track) and
Montenegrin as a 5-task preview:… See the full description on the dataset page: https://huggingface.co/datasets/permitt/superglue-hr.
