CoolFace
Datasetpublic

pbhappliedsystems/quant_eval_run_provenance

quant_eval — Run provenance One row per published run: model identity, contract identifiers, fixture hash, decoding conditions, licence, and the SHA-256 and byte size of both weight artifacts. Accompanied by the calibration lineage that informed each published run. Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing. Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_run_provenance.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes55downloads
Dataset Card

quant_eval — Run provenance

One row per published run: model identity, contract identifiers, fixture hash, decoding conditions, licence, and the SHA-256 and byte size of both weight artifacts. Accompanied by the calibration lineage that informed each published run.

Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.

Cite this dataset: 10.5281/zenodo.22010462 — concept DOI, always resolves to the latest version. This exact deposit: 10.5281/zenodo.22010463 — version DOI, frozen. Cite this one where reported numbers must stay verifiable against the object referenced.

What this file contains

FileRowsColumns
quant_eval_run_provenance.csv641
calibration_lineage.csv144

Supporting files: source_bundle_checksums.json.

Corpus scope

RunModelBaselineQuantizedSubstrateLicence
Mistral_Nemo_Instruct_2407_20260814_030505mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq4k_mlocalApache-2.0
Mistral_Nemo_Instruct_2407_20260815_113254mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq5k_mlocalApache-2.0
Mistral_Nemo_Instruct_2407_20260816_084553mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq80localApache-2.0
Qwen2.5_14B_Instruct_1M_20260815_220633Qwen/Qwen2.5-14B-Instruct-1Mmodal_f16modalq4k_mModalApache-2.0
Qwen2.5_32B_Instruct_20260815_081051Qwen/Qwen2.5-32B-Instructmodal_f16modalq4k_mModalApache-2.0
Qwen2.5_7B_Instruct_20260814_234822Qwen/Qwen2.5-7B-Instructgguf_f16ggufq4k_mlocalApache-2.0

Every run evaluates a full-weight baseline and a quantized variant of the same model against the identical locked fixture set, case for case. Statistical comparison is paired: the two-sided exact McNemar test on per-case outcomes, with Wilson intervals on the rates.

Columns

quant_eval_run_provenance.csv

  run_id                                        model_id                                      canonical_upstream_model_id
  adapter_id                                    adapter_reason                                version_tag
  evaluation_contract_id                        scoring_contract                              prompt_contract
  fixture_construction_contract                 fixtures_sha256                               fixture_version_label
  fixture_split                                 profile                                       run_purpose
  promotion_status                              publication_status                            release_validation_status
  seed                                          fixture_generation_seed                       timestamp
  baseline_quant_type                           quantized_quant_type                          baseline_runner
  quantized_runner                              execution_substrate                           decode_temperature
  decode_seed_status                            decode_context_size                           decode_top_p
  decode_top_k                                  license                                       spdx_license_id
  commercial_use_status                         baseline_artifact_sha256                      baseline_artifact_bytes
  quantized_artifact_sha256                     quantized_artifact_bytes                      upstream_revision
  upstream_revision_status                      rows

calibration_lineage.csv

  published_run_id                              calibration_run_id                            role
  published

Verification

This corpus is derived from sanitized publication bundles produced by the quant_eval harness. It is designed to be checked rather than trusted:

  • source_bundle_checksums.json, included here, republishes, verbatim, the SHA-256 digest and byte length of every file in every source bundle. No source file was modified.
  • Before this file was written, the builder verified all 72 source-file digests and independently recomputed all 96 family x runner pass rates from the raw per-case rows, matching the harness rollups exactly.
  • The fields here are run-level metadata carried through from the source bundles unchanged, not statistics derived from per-case rows. They are traceable through the bundle digests above. Recomputable aggregates live in the paired degradation statistics (D5) and family pass rates (D6) datasets, both derived from the per-case results dataset (D1).

Limits you should know before using this

  • Decoding conditions are not uniform across models. Temperature follows each publisher's own model card, so cross-model comparison of absolute pass rates is confounded. Within-run pairing is unaffected, which is what the paired test requires. The conditions are published per row and per run so they can be filtered on.
  • Runs on the Modal substrate record `seed` status `unsupported` — the deployed method signature accepts no seed — so those runs are not exactly reproducible. Local runs applied a fixed seed.
  • Runtime figures are observed harness wall time on the recorded hardware and backends. They are not a controlled throughput benchmark and not a general claim about quantization performance at any precision on any hardware. Direction is published as an explicit label because not every measured pair is a speedup.
  • The fuzz family is an adaptive trajectory evaluated from identical starting fixtures. Its paired test compares complete case outcomes, not identical post-divergence prompts.
  • `upstream_revision` is empty for 2 of 6 runs (Qwen2.5_14B_Instruct_1M_20260815_220633, Qwen2.5_7B_Instruct_20260814_234822). Where present, it is the exact upstream repository revision the evaluated weights were converted from. Where empty, upstream_revision_status records why — in this corpus, unavailable_for_pre_provenance_legacy_artifact. This is a documented absence, not a dropped measurement: weight identity for those runs is still established exactly, by the artifact SHA-256 recorded in this file, but not tied to a named upstream commit.
  • Calibration runs are not published. Runs that informed a published run are disclosed by identifier in calibration_lineage.csv, included here, so the record is complete without releasing provisional numbers.

Citation

bibtex
@dataset{pbh_quant_eval_d4,
  author    = {Hill, Patrick},
  title     = {quant_eval Run provenance},
  publisher = {PBH Applied Systems, LLC},
  year      = {2026},
  doi       = {10.5281/zenodo.22010462},
  note      = {Version DOI: 10.5281/zenodo.22010463},
  license   = {CC-BY-4.0}
}

Licence

Creative Commons Attribution 4.0 International (CC BY 4.0). See LICENSE. Commercial use is permitted; attribution is required.

This corpus describes third-party models and redistributes no model weights. Each evaluated model remains under its own licence, recorded per run in the run provenance dataset.


Produced by builddatasets.py 2.4.0 from quanteval publication bundles. Built 2026-08-19.