datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quant_eval_run_provenance
quant_eval — Run provenance
One row per published run: model identity, contract identifiers, fixture hash, decoding conditions, licence, and the SHA-256 and byte size of both weight artifacts. Accompanied by the calibration lineage that informed each published run.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_run_provenance.quant_eval_v7_21_per_case_results_and_run_provenance
quant_eval v7.21 — Per-Case Evaluation Results and Run Provenance
Supplementary evidence for the whitepaper quant_eval: A Behavioral Evaluation Harness for
Full-Weight and Quantized Large Language Models.
Author: Patrick Hill, PBH Applied Systems, LLC
ORCID: 0009-0008-3662-1681
Licence: CC BY 4.0
Concept DOI (all versions): 10.5281/zenodo.22851375
Version DOI (this deposit): 10.5281/zenodo.22851376
What this deposit is
Every quantitative result reported in the… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_v7_21_per_case_results_and_run_provenance.
