CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NuBerea /authority-provenancegated authority-provenance A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it records independent signals bearing on the authority of the text at that point: textual stability (is the reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?), and canonical reception (how the church received it). These axes are kept separate so that questions of manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.textfeature-extraction10K<n<100K0 likes452 downloads24d agoHugging Face02csoai /gspc-provenance-controls GSPC — provenance controls facts (ChainFacts) SWIFT census (live): https://councilof.ai/api/swift XRPL reader (live): https://councilof.ai/api/xrpl MEASURED financial/domain axis (issuer-account / on-chain control facts, n=6). Not a model leaderboard. No accuracy, no fleet, no leader, no separation — measured is not scored. Frozen bank on Hub. Live n and status are the provenance-controls row on GET https://councilof.ai/api/gspc. Not a certificate. Council of AI · CSOAI Ltd… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-provenance-controls.tabularothern<1K0 likes387 downloads8d agoHugging Face03trucyberlab /multimodal-ICS-provenance ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/trucyberlab/multimodal-ICS-provenance.tabulargraph-ml100K<n<1M0 likes233 downloads3mo agoHugging Face04siddharthmb /mats-gf-provenance-corpora Provenance-codeword training corpora All training corpora from the eight-experiment provenance codewords program (per-source activation codewords in Qwen3 models). Code, paper, and reproduction scripts: https://github.com/Sid-MB/mats-gf-provenance-codewords Each synthetic corpus ships in full: docs.parquet (training documents), train.parquet, qa.parquet (probe questions incl. phantom-fact controls), generation intermediates (raw/), the sqlite sequence store (seqdb/), and audit… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-corpora.tabular1K<n<10K0 likes231 downloads2mo agoHugging Face05eda-bench-neurips-2026 /eda-bench-provenance EDA Bench Evaluation Provenance This dataset contains published evaluation records for EDA Bench. Use it to audit reported runs and inspect the recorded execution environment. Contents The repository includes records such as: harness and tool version metadata task prompts and runtime task copies model responses and command records standard output and error logs raw and normalized grading metrics run-level reports source snapshots used to identify the evaluated… See the full description on the dataset page: https://huggingface.co/datasets/eda-bench-neurips-2026/eda-bench-provenance.text0 likes173 downloads26d agoHugging Face06leesharks /provenance-erasureall things are now lawful to you in jack feist EA-RHIZOME-PER-01 — provenance erasure Where provenance disappears, and what measures it. A transformation chain — source, retrieval, selection, composition, summarisation, reception — with the instruments that score retention at each step, and the cases where the instruments were applied to themselves. These are symbola. They are for traversal. A token broken in two, each half held by a different party, no half… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/provenance-erasure.textn<1K4 likes169 downloads10h agoHugging Face07common-pile /data_provenance_initiative_filtered Data Provenance Initiative Description The Data Provenance Initiative is a digital library of supervised datasets that have been manually annotated with their source and license information [ 104, 107 ]. We leverage their tooling to filter HuggingFace datasets, based on a range of criteria, including their licenses. Specifically, we filter the data according to these criteria: contains English language or code data, the text is not model-generated, the dataset’s audit… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/data_provenance_initiative_filtered.texttext-generation1M<n<10M0 likes168 downloads1y agoHugging Face08mjbommar /opengloss-v2.2-provenance Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility. OpenGloss v2.2 — Provenance The audit trail for OpenGloss v2.2: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-provenance.tabulartext-classification1M<n<10M0 likes155 downloads13d agoHugging Face09oof-baroomf /eda-bench-provenancetext1 likes154 downloads5mo agoHugging Face10cisco-ai /model-provenance-kit Deep-Signal Weight Fingerprints Dataset Summary This repository distributes pre-computed deep-signal weight fingerprints for use with Model Provenance Kit (Model ProvenanceKit). The primary deliverable is deep-signals.zip: a compressed archive of Apache Parquet files that store dense numerical features extracted from publicly released transformer (and related) model weights on the Hugging Face Hub or equivalent sources. Each file encodes multi-signal weight-level… See the full description on the dataset page: https://huggingface.co/datasets/cisco-ai/model-provenance-kit.other1K<n<10K4 likes132 downloads2mo agoHugging Face11mjbommar /opengloss-v2.3-provenance OpenGloss v2.3 — Provenance The audit trail for OpenGloss v2.3: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the prompt hit the provider's cache, and what it cost. Nothing in this release was written without a row here. It is what makes the cost claims in the other cards checkable rather than asserted, and it is what a reader who wants to know which model wrote this field should join… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-provenance.tabulartext-classification1M<n<10M0 likes127 downloads13d agoHugging Face12mjbommar /opengloss-v2.0-provenance Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Provenance The audit trail for OpenGloss v2.0: one row per recorded unit of work, saying which stage ran, which model answered, how many prompt and completion tokens it used, how much of the prompt hit the provider's cache, and what it cost.… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-provenance.tabulartext-classification1M<n<10M0 likes125 downloads15d agoHugging Face13mjbommar /opengloss-v2.1-provenance Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Provenance The audit trail for OpenGloss v2.1: one row per recorded unit of work, saying which stage ran, which model… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-provenance.tabulartext-classification1M<n<10M0 likes116 downloads13d agoHugging Face14common-pile /data_provenance_initiative Data Provenance Initiative Description The Data Provenance Initiative is a digital library of supervised datasets that have been manually annotated with their source and license information [ 104, 107 ]. We leverage their tooling to filter HuggingFace datasets, based on a range of criteria, including their licenses. Specifically, we filter the data according to these criteria: contains English language or code data, the text is not model-generated, the dataset’s audit… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/data_provenance_initiative.text1M<n<10M0 likes87 downloads1y agoHugging Face15Lexsi /provenance-grounded-synthetic-qa synthetic_qa_data This dataset contains synthetic question-answer pairs generated and filtered using the following models: Generation Models Qwen/Qwen3-1.7B Qwen/Qwen3-4B Qwen/Qwen3-8B Filtering Model Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters Dataset Structure data/ ├── unfiltered_qa/ # Raw generated QA pairs per model ├── both_filtered_qa/ # QA pairs passing both filters ├──… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/provenance-grounded-synthetic-qa.text10K<n<100K0 likes75 downloads3mo agoHugging Face16oof-baroomf /cad-bench-provenance0 likes74 downloads5mo agoHugging Face17sciexp /insdc-provenance0 likes74 downloads3mo agoHugging Face18siddharthmb /mats-gf-provenance-readouts Provenance-codeword readout records Per-experiment readout artifacts from the provenance codewords program: source decodes, generation-time attribution records, model answers, suppression/ablation sweeps, calibration files, capability checks (MMLU / perplexity), and comparison matrices — the evidence behind every number in the paper. JSON/parquet only; contents are synthetic or model-generated text plus scores (no raw news text, no activation tensors — those stay on-cluster and… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-provenance-readouts.tabularn<1K0 likes73 downloads2mo agoHugging Face19harryCJ /multimodal-ICS-provenance ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.tabulargraph-ml100K<n<1M0 likes64 downloads2mo agoHugging Face20Contrastive-Prov /Provenance0 likes63 downloads2y agoHugging Face21arpandeepk /code-atlas-provenance Code-ATLAS Corpus Registry This repository is the data-first foundation for adapting ATLAS to programming languages. It is intentionally a registry and split specification before it is a training corpus. The scientific goal is to measure directed transfer among roughly 20 programming languages, fit loss-based scaling laws, and predict a data mixture for adapting a model to a low-resource or newly introduced language. Publication contract Every published payload… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/code-atlas-provenance.text-generation0 likes62 downloads1mo agoHugging Face22CAD-bench /cad-eda-public-provenance CAD/EDA Benchmark Source Provenance This dataset records the public source revisions and license identifiers used to derive selected electronic-design benchmark tasks. It contains provenance metadata only. It does not redistribute source files, benchmark answers, customer material, model outputs, credentials, or personal data. Use each source under the license named in its record. The source repository remains the authority for its license text and revision history. textn<1K0 likes57 downloads26d agoHugging Face23pbhappliedsystems /quant_eval_run_provenance quant_eval — Run provenance One row per published run: model identity, contract identifiers, fixture hash, decoding conditions, licence, and the SHA-256 and byte size of both weight artifacts. Accompanied by the calibration lineage that informed each published run. Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing. Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_run_provenance.tabularn<1K0 likes55 downloads1mo agoHugging Face24thoughtworks /cbd-100pair-hf-natural-rewrite-provenance cbd-100pair-hf-natural-rewrite-provenance Provenance-complete regeneration of 500 hf_natural-style conjunctive-pair rewrites: five rows for each of the 100 frozen AND-trigger pairs. This is a new deterministic sample, not a recovery of the historical 14,961 rows. The old published data removed source_id, so its exact OpenOrca-to-rewrite mapping cannot be reconstructed. Each source is an indexed row from Open-Orca/OpenOrca. Rewrites were generated with the same model and prompt… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-100pair-hf-natural-rewrite-provenance.tabulartext-generationn<1K0 likes54 downloads4d agoHugging Face25Om22s /alexandria-attribution-provenance-20260913 Alexandria speaker-attribution evaluation provenance This release contains aggregate, paired evaluation measurements for a Qwen3-14B speaker-attribution LoRA. It is a reproducibility record, not a training or evaluation corpus. Contents results.json contains aggregate counts, accuracy, strict shared-row counts, decoding settings, and evaluator commit. provenance.json maps each aggregate record to the SHA-256 of its private source artifact and to hashes of the… See the full description on the dataset page: https://huggingface.co/datasets/Om22s/alexandria-attribution-provenance-20260913.text-classification0 likes49 downloads1d agoHugging Face26rafmacalaba /data-use-provenance-sft Data-use provenance SFT Instruction-following examples for extracting provenance attributes (producer, year, geography, acronym) of a data mention from its context. Labels are generated by gpt-5.6-luna (batch API) and verbatim-filtered. Format ChatML messages: a user prompt (mention + context) and an assistant JSON answer. Absent attributes are omitted. {"train": 16274, "val": 3488, "holdout": 3487} text10K<n<100K0 likes43 downloads1mo agoHugging Face27eda-bench-neurips-2026 /eda-bench-public-provenance EDA Benchmark Public Provenance This dataset records aggregate provenance for the privacy-redacted EDA benchmark extension. It contains no payload files, credentials, personal data, browser data, customer material, or benchmark answers. The raw benchmark evidence remains private while its release rights are reviewed. tabularn<1K0 likes43 downloads16d agoHugging Face28dougdotcon /douvras-dataset-quality-provenance Douvras Dataset Quality and Provenance v0.1 Synthetic quality-gate records covering PII, duplicate rate, license, schema, split overlap and provenance. Labels are PASS, REVIEW and BLOCK; PII or cross-split overlap always blocks. It contains 36 records (24/6/6) across 12 artifact instances, split by artifact. This is a pre-publication diagnostic protocol. It contains no real datasets and never publishes automatically. textn<1K0 likes43 downloads9d agoHugging Face29eda-bench-neurips-2026 /eda-bench-raw-provenance EDA Bench raw-extension provenance This public record documents a private EDA Bench raw extension. It contains no raw designs, source files, account data, or personal information. The private extension contains 263,257 payload files totaling 61,558,050,574 bytes. Its immutable inclusion manifest, privacy receipt, archive, remote restore, and source-to-restore Git executable-bit parity were verified before this record was prepared. provenance.json contains content hashes and… See the full description on the dataset page: https://huggingface.co/datasets/eda-bench-neurips-2026/eda-bench-raw-provenance.textn<1K0 likes42 downloads15d agoHugging Face30liaolw /physx-v18-2-raw-part-names-provenance-corrected-20260912 PhysX v18.2 Raw Part-Name Release This public, label-only release contains the frozen 233-train / 64-validation cohort and its 297 selected original finaljson annotations. The part target for every record is exactly finaljson.parts[].name.strip(). The dataset does not include images, checkpoints, model files, Flow Matching / PhysX-3D data, renders, logs, or other training artifacts. Image rebasing train.jsonl and val.jsonl contain only portable relative image… See the full description on the dataset page: https://huggingface.co/datasets/liaolw/physx-v18-2-raw-part-names-provenance-corrected-20260912.1 likes34 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.