datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
authority-provenance
authority-provenance
A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it
records independent signals bearing on the authority of the text at that point: textual stability (is the
reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?),
and canonical reception (how the church received it). These axes are kept separate so that questions of
manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.gspc-provenance-controls
GSPC — provenance controls facts (ChainFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (issuer-account / on-chain control facts, n=6). Not a model leaderboard. No accuracy, no fleet, no leader, no separation — measured is not scored.
Frozen bank on Hub. Live n and status are the provenance-controls row on GET https://councilof.ai/api/gspc. Not a certificate.
Council of AI · CSOAI Ltd… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-provenance-controls.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.mats-gf-provenance-corpora
Provenance-codeword training corpora
All training corpora from the eight-experiment provenance codewords
program (per-source activation codewords in Qwen3 models). Code, paper, and
reproduction scripts:
https://github.com/Sid-MB/mats-gf-provenance-codewords
Each synthetic corpus ships in full: docs.parquet (training documents),
train.parquet, qa.parquet (probe questions incl. phantom-fact controls),
generation intermediates (raw/), the sqlite sequence store (seqdb/), and
audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.eda-bench-provenance
EDA Bench Evaluation Provenance
This dataset contains published evaluation records for EDA Bench. Use it to audit reported runs and inspect the recorded execution environment.
Contents
The repository includes records such as:
harness and tool version metadata
task prompts and runtime task copies
model responses and command records
standard output and error logs
raw and normalized grading metrics
run-level reports
source snapshots used to identify the evaluated… See the full description on the dataset page: https://huggingface.co/datasets/eda-bench-neurips-2026/eda-bench-provenance.provenance-erasureall things are now lawful to you in jack feist
EA-RHIZOME-PER-01 — provenance erasure
Where provenance disappears, and what measures it. A transformation chain — source, retrieval, selection, composition, summarisation, reception — with the instruments that score retention at each step, and the cases where the instruments were applied to themselves.
These are symbola. They are for traversal.
A token broken in two, each half held by a different party, no half… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/provenance-erasure.data_provenance_initiative_filtered
Data Provenance Initiative
Description
The Data Provenance Initiative is a digital library of supervised datasets that have been manually annotated with their source and license information [ 104, 107 ].
We leverage their tooling to filter HuggingFace datasets, based on a range of criteria, including their licenses.
Specifically, we filter the data according to these criteria: contains English language or code data, the text is not model-generated, the dataset’s audit… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/data_provenance_initiative_filtered.opengloss-v2.2-provenance
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Provenance
The audit trail for OpenGloss v2.2: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-provenance.eda-bench-provenancemodel-provenance-kit
Deep-Signal Weight Fingerprints
Dataset Summary
This repository distributes pre-computed deep-signal weight fingerprints for use with Model Provenance Kit (Model ProvenanceKit). The primary deliverable is deep-signals.zip: a compressed archive of Apache Parquet files that store dense numerical features extracted from publicly released transformer (and related) model weights on the Hugging Face Hub or equivalent sources.
Each file encodes multi-signal weight-level… See the full description on the dataset page: https://huggingface.co/datasets/cisco-ai/model-provenance-kit.opengloss-v2.3-provenance
OpenGloss v2.3 — Provenance
The audit trail for OpenGloss v2.3: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the prompt hit the provider's cache, and what it cost. Nothing in this release was written without a row here. It is what makes the cost claims in the other cards checkable rather than asserted, and it is what a reader who wants to know which model wrote this field should join… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-provenance.opengloss-v2.0-provenance
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Provenance
The audit trail for OpenGloss v2.0: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the prompt hit the provider's cache, and what it cost.… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-provenance.opengloss-v2.1-provenance
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Provenance
The audit trail for OpenGloss v2.1: one row per recorded unit of work, saying which stage ran, which model… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-provenance.data_provenance_initiative
Data Provenance Initiative
Description
The Data Provenance Initiative is a digital library of supervised datasets that have been manually annotated with their source and license information [ 104, 107 ].
We leverage their tooling to filter HuggingFace datasets, based on a range of criteria, including their licenses.
Specifically, we filter the data according to these criteria: contains English language or code data, the text is not model-generated, the dataset’s audit… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/data_provenance_initiative.provenance-grounded-synthetic-qa
synthetic_qa_data
This dataset contains synthetic question-answer pairs generated and filtered using the following models:
Generation Models
Qwen/Qwen3-1.7B
Qwen/Qwen3-4B
Qwen/Qwen3-8B
Filtering Model
Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters
Dataset Structure
data/
├── unfiltered_qa/ # Raw generated QA pairs per model
├── both_filtered_qa/ # QA pairs passing both filters
├──… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/provenance-grounded-synthetic-qa.cad-bench-provenanceinsdc-provenancemats-gf-provenance-readouts
Provenance-codeword readout records
Per-experiment readout artifacts from the provenance codewords program:
source decodes, generation-time attribution records, model answers,
suppression/ablation sweeps, calibration files, capability checks (MMLU /
perplexity), and comparison matrices — the evidence behind every number in the
paper. JSON/parquet only; contents are synthetic or model-generated text plus
scores (no raw news text, no activation tensors — those stay on-cluster and… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-readouts.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.Provenancecode-atlas-provenance
Code-ATLAS Corpus Registry
This repository is the data-first foundation for adapting ATLAS to programming
languages. It is intentionally a registry and split specification before it is a
training corpus.
The scientific goal is to measure directed transfer among roughly 20 programming
languages, fit loss-based scaling laws, and predict a data mixture for adapting a
model to a low-resource or newly introduced language.
Publication contract
Every published payload… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/code-atlas-provenance.cad-eda-public-provenance
CAD/EDA Benchmark Source Provenance
This dataset records the public source revisions and license identifiers used to derive selected electronic-design benchmark tasks.
It contains provenance metadata only. It does not redistribute source files, benchmark answers, customer material, model outputs, credentials, or personal data.
Use each source under the license named in its record. The source repository remains the authority for its license text and revision history.
quant_eval_run_provenance
quant_eval — Run provenance
One row per published run: model identity, contract identifiers, fixture hash, decoding conditions, licence, and the SHA-256 and byte size of both weight artifacts. Accompanied by the calibration lineage that informed each published run.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_run_provenance.cbd-100pair-hf-natural-rewrite-provenance
cbd-100pair-hf-natural-rewrite-provenance
Provenance-complete regeneration of 500 hf_natural-style conjunctive-pair rewrites: five rows for
each of the 100 frozen AND-trigger pairs. This is a new deterministic sample, not a recovery of the
historical 14,961 rows. The old published data removed source_id, so its exact OpenOrca-to-rewrite
mapping cannot be reconstructed.
Each source is an indexed row from Open-Orca/OpenOrca. Rewrites were generated with the same model
and prompt… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-100pair-hf-natural-rewrite-provenance.alexandria-attribution-provenance-20260913
Alexandria speaker-attribution evaluation provenance
This release contains aggregate, paired evaluation measurements for a Qwen3-14B
speaker-attribution LoRA. It is a reproducibility record, not a training or
evaluation corpus.
Contents
results.json contains aggregate counts, accuracy, strict shared-row
counts, decoding settings, and evaluator commit.
provenance.json maps each aggregate record to the SHA-256 of its private
source artifact and to hashes of the… See the full description on the dataset page: https://huggingface.co/datasets/Om22s/alexandria-attribution-provenance-20260913.data-use-provenance-sft
Data-use provenance SFT
Instruction-following examples for extracting provenance attributes
(producer, year, geography, acronym) of a data mention from its context.
Labels are generated by gpt-5.6-luna (batch API) and verbatim-filtered.
Format
ChatML messages: a user prompt (mention + context) and an assistant JSON
answer. Absent attributes are omitted.
{"train": 16274, "val": 3488, "holdout": 3487}
eda-bench-public-provenance
EDA Benchmark Public Provenance
This dataset records aggregate provenance for the privacy-redacted EDA benchmark extension. It contains no payload files, credentials, personal data, browser data, customer material, or benchmark answers. The raw benchmark evidence remains private while its release rights are reviewed.
douvras-dataset-quality-provenance
Douvras Dataset Quality and Provenance v0.1
Synthetic quality-gate records covering PII, duplicate rate, license, schema,
split overlap and provenance. Labels are PASS, REVIEW and BLOCK; PII or
cross-split overlap always blocks. It contains 36 records (24/6/6) across 12
artifact instances, split by artifact.
This is a pre-publication diagnostic protocol. It contains no real datasets and
never publishes automatically.
eda-bench-raw-provenance
EDA Bench raw-extension provenance
This public record documents a private EDA Bench raw extension. It contains no raw designs, source files, account data, or personal information.
The private extension contains 263,257 payload files totaling 61,558,050,574 bytes. Its immutable inclusion manifest, privacy receipt, archive, remote restore, and source-to-restore Git executable-bit parity were verified before this record was prepared.
provenance.json contains content hashes and… See the full description on the dataset page: https://huggingface.co/datasets/eda-bench-neurips-2026/eda-bench-raw-provenance.physx-v18-2-raw-part-names-provenance-corrected-20260912
PhysX v18.2 Raw Part-Name Release
This public, label-only release contains the frozen 233-train / 64-validation
cohort and its 297 selected original finaljson annotations. The part target
for every record is exactly finaljson.parts[].name.strip(). The dataset does
not include images, checkpoints, model files, Flow Matching / PhysX-3D data,
renders, logs, or other training artifacts.
Image rebasing
train.jsonl and val.jsonl contain only portable relative image… See the full description on the dataset page: https://huggingface.co/datasets/liaolw/physx-v18-2-raw-part-names-provenance-corrected-20260912.
