datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-provenance-controls
GSPC — provenance controls facts (ChainFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (issuer-account / on-chain control facts, n=6). Not a model leaderboard. No accuracy, no fleet, no leader, no separation — measured is not scored.
Frozen bank on Hub. Live n and status are the provenance-controls row on GET https://councilof.ai/api/gspc. Not a certificate.
Council of AI · CSOAI Ltd… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-provenance-controls.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.mats-gf-provenance-corpora
Provenance-codeword training corpora
All training corpora from the eight-experiment provenance codewords
program (per-source activation codewords in Qwen3 models). Code, paper, and
reproduction scripts:
https://github.com/Sid-MB/mats-gf-provenance-codewords
Each synthetic corpus ships in full: docs.parquet (training documents),
train.parquet, qa.parquet (probe questions incl. phantom-fact controls),
generation intermediates (raw/), the sqlite sequence store (seqdb/), and
audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.opengloss-v2.2-provenance
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Provenance
The audit trail for OpenGloss v2.2: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-provenance.opengloss-v2.3-provenance
OpenGloss v2.3 — Provenance
The audit trail for OpenGloss v2.3: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the prompt hit the provider's cache, and what it cost. Nothing in this release was written without a row here. It is what makes the cost claims in the other cards checkable rather than asserted, and it is what a reader who wants to know which model wrote this field should join… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-provenance.opengloss-v2.0-provenance
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Provenance
The audit trail for OpenGloss v2.0: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the prompt hit the provider's cache, and what it cost.… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-provenance.opengloss-v2.1-provenance
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Provenance
The audit trail for OpenGloss v2.1: one row per recorded unit of work, saying which stage ran, which model… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-provenance.mats-gf-provenance-readouts
Provenance-codeword readout records
Per-experiment readout artifacts from the provenance codewords program:
source decodes, generation-time attribution records, model answers,
suppression/ablation sweeps, calibration files, capability checks (MMLU /
perplexity), and comparison matrices — the evidence behind every number in the
paper. JSON/parquet only; contents are synthetic or model-generated text plus
scores (no raw news text, no activation tensors — those stay on-cluster and… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-readouts.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.cbd-100pair-hf-natural-rewrite-provenance
cbd-100pair-hf-natural-rewrite-provenance
Provenance-complete regeneration of 500 hf_natural-style conjunctive-pair rewrites: five rows for
each of the 100 frozen AND-trigger pairs. This is a new deterministic sample, not a recovery of the
historical 14,961 rows. The old published data removed source_id, so its exact OpenOrca-to-rewrite
mapping cannot be reconstructed.
Each source is an indexed row from Open-Orca/OpenOrca. Rewrites were generated with the same model
and prompt… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-100pair-hf-natural-rewrite-provenance.quant_eval_run_provenance
quant_eval — Run provenance
One row per published run: model identity, contract identifiers, fixture hash, decoding conditions, licence, and the SHA-256 and byte size of both weight artifacts. Accompanied by the calibration lineage that informed each published run.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_run_provenance.eda-bench-public-provenance
EDA Benchmark Public Provenance
This dataset records aggregate provenance for the privacy-redacted EDA benchmark extension. It contains no payload files, credentials, personal data, browser data, customer material, or benchmark answers. The raw benchmark evidence remains private while its release rights are reviewed.
condition-provenance-40
Condition provenance in scientific claim verification (40 claims)
Summary
Forty claims about NMC811 cathodes, each paired with one open-access source paper. A model decides whether the paper's measurements match every condition in the claim, and answers supported, not_supported, or no_comparable_evidence. In 12 of the 40 claims the paper never states the experimental condition the claim turns on, so the keyed answer is no_comparable_evidence, while the other 28… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/condition-provenance-40.quant_eval_v7_21_per_case_results_and_run_provenance
quant_eval v7.21 — Per-Case Evaluation Results and Run Provenance
Supplementary evidence for the whitepaper quant_eval: A Behavioral Evaluation Harness for
Full-Weight and Quantized Large Language Models.
Author: Patrick Hill, PBH Applied Systems, LLC
ORCID: 0009-0008-3662-1681
Licence: CC BY 4.0
Concept DOI (all versions): 10.5281/zenodo.22851375
Version DOI (this deposit): 10.5281/zenodo.22851376
What this deposit is
Every quantitative result reported in the… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_v7_21_per_case_results_and_run_provenance.verifiable-ai-provenance-bench
Verifiable AI Provenance Bench (TTTPS)
25 real timestamp-provenance receipts generated on 2026-08-04 by calling the
live KPP (Kenosian Protocol Platform) provenance API
(POST /v1/anchor, POST /v1/verify), which implements the TTTPS (Time-Token
Tamper-evident Provenance Seal) scheme. Each row is one real API round trip:
a content_hash was submitted to /v1/anchor, the returned receipt_id was then
submitted to /v1/verify, and both raw responses are recorded.
This dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/Pittro/verifiable-ai-provenance-bench.crcs-provenance-trail
