pbhappliedsystems/quant_eval_throughput_telemetry
quant_eval — Throughput telemetry One record per generation call across all published runs, pooled into a single union schema. Carries token counts, timing, and the token-count source per call. Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing. Cite this dataset: 10.5281/zenodo.22009987 — concept DOI, always resolves to the latest… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_throughput_telemetry.
quant_eval — Throughput telemetry
One record per generation call across all published runs, pooled into a single union schema. Carries token counts, timing, and the token-count source per call.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset: 10.5281/zenodo.22009987 — concept DOI, always resolves to the latest version. This exact deposit: 10.5281/zenodo.22009988 — version DOI, frozen. Cite this one where reported numbers must stay verifiable against the object referenced.
What this file contains
Supporting files: data_dictionary.json, source_bundle_checksums.json.
Corpus scope
Every run evaluates a full-weight baseline and a quantized variant of the same model against the identical locked fixture set, case for case. Statistical comparison is paired: the two-sided exact McNemar test on per-case outcomes, with Wilson intervals on the rates.
Columns
quant_eval_throughput_telemetry.jsonl
attempt call call_id
case_id cumulative_generated_tokens cumulative_generation_seconds
cumulative_prompt_tokens cumulative_tokens_per_second evaluation_contract_id
event family finish_reason
fixtures_sha256 generated_tokens generation_seconds
max_new_tokens max_new_tokens_recorded model_id
prompt_tokens quant_type run_elapsed_seconds
run_id runner schema_version
scored stage timestamp_unix
token_count_source token_limit_reached tokens_per_second
version_tagdata_dictionary.json, included here, documents every field: applicability, data type, evidence role, gate membership, the meaning of an empty value, and observed population counts. An empty cell is not automatically a missing measurement — several fields are conditional on task family, and their empty-value meaning is recorded explicitly.
Verification
This corpus is derived from sanitized publication bundles produced by the quant_eval harness. It is designed to be checked rather than trusted:
source_bundle_checksums.json, included here, republishes, verbatim, the SHA-256 digest and byte length of every file in every source bundle. No source file was modified.- Before this file was written, the builder verified all 72 source-file digests and independently recomputed all 96 family x runner pass rates from the raw per-case rows, matching the harness rollups exactly.
- Aggregates published elsewhere in this corpus — the paired degradation statistics (D5) and family pass rates (D6) datasets — are recomputable from the per-case results dataset (D1) using the gate definitions in
data_dictionary.json, included here.
Limits you should know before using this
- Decoding conditions are not uniform across models. Temperature follows each publisher's own model card, so cross-model comparison of absolute pass rates is confounded. Within-run pairing is unaffected, which is what the paired test requires. The conditions are published per row and per run so they can be filtered on.
- Runs on the Modal substrate record `seed` status `unsupported` — the deployed method signature accepts no seed — so those runs are not exactly reproducible. Local runs applied a fixed seed.
- Runtime figures are observed harness wall time on the recorded hardware and backends. They are not a controlled throughput benchmark and not a general claim about quantization performance at any precision on any hardware. Direction is published as an explicit label because not every measured pair is a speedup.
- The fuzz family is an adaptive trajectory evaluated from identical starting fixtures. Its paired test compares complete case outcomes, not identical post-divergence prompts.
- Calibration runs are not published. Runs that informed a published run are disclosed by identifier in
calibration_lineage.csv, published in the run provenance dataset (D4), so the record is complete without releasing provisional numbers.
Citation
@dataset{pbh_quant_eval_d2,
author = {Hill, Patrick},
title = {quant_eval Throughput telemetry},
publisher = {PBH Applied Systems, LLC},
year = {2026},
doi = {10.5281/zenodo.22009987},
note = {Version DOI: 10.5281/zenodo.22009988},
license = {CC-BY-4.0}
}Licence
Creative Commons Attribution 4.0 International (CC BY 4.0). See LICENSE. Commercial use is permitted; attribution is required.
This corpus describes third-party models and redistributes no model weights. Each evaluated model remains under its own licence, recorded per run in the run provenance dataset.
Produced by builddatasets.py 2.4.0 from quanteval publication bundles. Built 2026-08-19.
