datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-provenance-controls
GSPC — provenance controls facts (ChainFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (issuer-account / on-chain control facts, n=6). Not a model leaderboard. No accuracy, no fleet, no leader, no separation — measured is not scored.
Frozen bank on Hub. Live n and status are the provenance-controls row on GET https://councilof.ai/api/gspc. Not a certificate.
Council of AI · CSOAI Ltd… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-provenance-controls.provenance-erasureall things are now lawful to you in jack feist
EA-RHIZOME-PER-01 — provenance erasure
Where provenance disappears, and what measures it. A transformation chain — source, retrieval, selection, composition, summarisation, reception — with the instruments that score retention at each step, and the cases where the instruments were applied to themselves.
These are symbola. They are for traversal.
A token broken in two, each half held by a different party, no half… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/provenance-erasure.data_provenance_initiative_filtered
Data Provenance Initiative
Description
The Data Provenance Initiative is a digital library of supervised datasets that have been manually annotated with their source and license information [ 104, 107 ].
We leverage their tooling to filter HuggingFace datasets, based on a range of criteria, including their licenses.
Specifically, we filter the data according to these criteria: contains English language or code data, the text is not model-generated, the dataset’s audit… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/data_provenance_initiative_filtered.data_provenance_initiative
Data Provenance Initiative
Description
The Data Provenance Initiative is a digital library of supervised datasets that have been manually annotated with their source and license information [ 104, 107 ].
We leverage their tooling to filter HuggingFace datasets, based on a range of criteria, including their licenses.
Specifically, we filter the data according to these criteria: contains English language or code data, the text is not model-generated, the dataset’s audit… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/data_provenance_initiative.provenance-grounded-synthetic-qa
synthetic_qa_data
This dataset contains synthetic question-answer pairs generated and filtered using the following models:
Generation Models
Qwen/Qwen3-1.7B
Qwen/Qwen3-4B
Qwen/Qwen3-8B
Filtering Model
Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters
Dataset Structure
data/
├── unfiltered_qa/ # Raw generated QA pairs per model
├── both_filtered_qa/ # QA pairs passing both filters
├──… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/provenance-grounded-synthetic-qa.cad-eda-public-provenance
CAD/EDA Benchmark Source Provenance
This dataset records the public source revisions and license identifiers used to derive selected electronic-design benchmark tasks.
It contains provenance metadata only. It does not redistribute source files, benchmark answers, customer material, model outputs, credentials, or personal data.
Use each source under the license named in its record. The source repository remains the authority for its license text and revision history.
douvras-dataset-quality-provenance
Douvras Dataset Quality and Provenance v0.1
Synthetic quality-gate records covering PII, duplicate rate, license, schema,
split overlap and provenance. Labels are PASS, REVIEW and BLOCK; PII or
cross-split overlap always blocks. It contains 36 records (24/6/6) across 12
artifact instances, split by artifact.
This is a pre-publication diagnostic protocol. It contains no real datasets and
never publishes automatically.
cbd-100pair-hf-natural-rewrite-provenance
cbd-100pair-hf-natural-rewrite-provenance
Provenance-complete regeneration of 500 hf_natural-style conjunctive-pair rewrites: five rows for
each of the 100 frozen AND-trigger pairs. This is a new deterministic sample, not a recovery of the
historical 14,961 rows. The old published data removed source_id, so its exact OpenOrca-to-rewrite
mapping cannot be reconstructed.
Each source is an indexed row from Open-Orca/OpenOrca. Rewrites were generated with the same model
and prompt… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-100pair-hf-natural-rewrite-provenance.eda-bench-raw-provenance
EDA Bench raw-extension provenance
This public record documents a private EDA Bench raw extension. It contains no raw designs, source files, account data, or personal information.
The private extension contains 263,257 payload files totaling 61,558,050,574 bytes. Its immutable inclusion manifest, privacy receipt, archive, remote restore, and source-to-restore Git executable-bit parity were verified before this record was prepared.
provenance.json contains content hashes and… See the full description on the dataset page: https://huggingface.co/datasets/eda-bench-neurips-2026/eda-bench-raw-provenance.eda-bench-public-provenance
EDA Benchmark Public Provenance
This dataset records aggregate provenance for the privacy-redacted EDA benchmark extension. It contains no payload files, credentials, personal data, browser data, customer material, or benchmark answers. The raw benchmark evidence remains private while its release rights are reviewed.
data-use-provenance-sft
Data-use provenance SFT
Instruction-following examples for extracting provenance attributes
(producer, year, geography, acronym) of a data mention from its context.
Labels are generated by gpt-5.6-luna (batch API) and verbatim-filtered.
Format
ChatML messages: a user prompt (mention + context) and an assistant JSON
answer. Absent attributes are omitted.
{"train": 16274, "val": 3488, "holdout": 3487}
github-actions-provenance-signed-untampered
GitHub Actions Provenance (Signed, Untampered)
This dataset contains signed provenance metadata in JSONL format generated from successful and untampered CI/CD builds using GitHub Actions. The data represents clean, valid examples of what software artifact provenance should look like in secure, uncompromised environments.
📂 Dataset Structure
Format: .jsonl files (each line is a JSON object)
Source: Generated by GitHub Actions CI/CD pipelines
Signature: Signed using… See the full description on the dataset page: https://huggingface.co/datasets/vchirrav/github-actions-provenance-signed-untampered.bitcoin-anchored-ai-provenance
Bitcoin-Anchored AI Provenance Receipts
Status (July 2026): this corpus is from the protocol's v1 "chain" era and is kept as a historical artifact. Every row remains independently verifiable: the OpenTimestamps proofs are portable and check against Bitcoin with the stock ots client, and the verify_url endpoint is still live. The protocol's current architecture is a C2SP transparency log cosigned by independent witnesses (log.markovianprotocol.com, browser verification at… See the full description on the dataset page: https://huggingface.co/datasets/MarkovianProtocol/bitcoin-anchored-ai-provenance.CurioCode-Provenance-Evidence-000573
CurioCode Provenance Evidence
This public dataset repository exposes machine-readable provenance observations used by the CurioCode release curation workflow. The records are evidence inputs, not a precomputed release ledger.
verifiable-ai-provenance-bench
Verifiable AI Provenance Bench (TTTPS)
25 real timestamp-provenance receipts generated on 2026-08-04 by calling the
live KPP (Kenosian Protocol Platform) provenance API
(POST /v1/anchor, POST /v1/verify), which implements the TTTPS (Time-Token
Tamper-evident Provenance Seal) scheme. Each row is one real API round trip:
a content_hash was submitted to /v1/anchor, the returned receipt_id was then
submitted to /v1/verify, and both raw responses are recorded.
This dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/Pittro/verifiable-ai-provenance-bench.han-data-provenance-tracking-v1
Humanoid Data Provenance Tracking Dataset
This dataset tracks the origin and transformation
of data used by humanoid AI systems.
It enhances transparency and security across the network.
Use Cases
Data integrity verification
Security auditing
Trust reinforcement
Fields
data_id
source_node
transformation_log
verification_status
Part of
Humanoid Network (HAN)
License
MIT
provenancebench
ProvenanceBench
A faithfulness + justified-abstention benchmark for regulated documentation. A correct
"I don't know" is scored as a first-class pass — credited only when the system abstains
for the right reason.
72 cases (cases config) over an 18-span synthetic regulated-style corpus
(corpus config) — a fictional QMS.
Each case is answerable (with gold supporting spans) or must-abstain (with a gold
reason from a published taxonomy: UAEval4RAG Table 6, cross-walked to… See the full description on the dataset page: https://huggingface.co/datasets/shryu1994/provenancebench.crcs-provenance-trail
