CoolFace
Datasetpublic

pbhappliedsystems/quant_eval_throughput_telemetry

quant_eval — Throughput telemetry One record per generation call across all published runs, pooled into a single union schema. Carries token counts, timing, and the token-count source per call. Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing. Cite this dataset: 10.5281/zenodo.22009987 — concept DOI, always resolves to the latest… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_throughput_telemetry.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes71downloads
Dataset Card

quant_eval — Throughput telemetry

One record per generation call across all published runs, pooled into a single union schema. Carries token counts, timing, and the token-count source per call.

Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.

Cite this dataset: 10.5281/zenodo.22009987 — concept DOI, always resolves to the latest version. This exact deposit: 10.5281/zenodo.22009988 — version DOI, frozen. Cite this one where reported numbers must stay verifiable against the object referenced.

What this file contains

FileRowsColumns
quant_eval_throughput_telemetry.jsonl27,37031

Supporting files: data_dictionary.json, source_bundle_checksums.json.

Corpus scope

RunModelBaselineQuantizedSubstrateLicence
Mistral_Nemo_Instruct_2407_20260814_030505mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq4k_mlocalApache-2.0
Mistral_Nemo_Instruct_2407_20260815_113254mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq5k_mlocalApache-2.0
Mistral_Nemo_Instruct_2407_20260816_084553mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq80localApache-2.0
Qwen2.5_14B_Instruct_1M_20260815_220633Qwen/Qwen2.5-14B-Instruct-1Mmodal_f16modalq4k_mModalApache-2.0
Qwen2.5_32B_Instruct_20260815_081051Qwen/Qwen2.5-32B-Instructmodal_f16modalq4k_mModalApache-2.0
Qwen2.5_7B_Instruct_20260814_234822Qwen/Qwen2.5-7B-Instructgguf_f16ggufq4k_mlocalApache-2.0

Every run evaluates a full-weight baseline and a quantized variant of the same model against the identical locked fixture set, case for case. Statistical comparison is paired: the two-sided exact McNemar test on per-case outcomes, with Wilson intervals on the rates.

Columns

quant_eval_throughput_telemetry.jsonl

  attempt                                       call                                          call_id
  case_id                                       cumulative_generated_tokens                   cumulative_generation_seconds
  cumulative_prompt_tokens                      cumulative_tokens_per_second                  evaluation_contract_id
  event                                         family                                        finish_reason
  fixtures_sha256                               generated_tokens                              generation_seconds
  max_new_tokens                                max_new_tokens_recorded                       model_id
  prompt_tokens                                 quant_type                                    run_elapsed_seconds
  run_id                                        runner                                        schema_version
  scored                                        stage                                         timestamp_unix
  token_count_source                            token_limit_reached                           tokens_per_second
  version_tag

data_dictionary.json, included here, documents every field: applicability, data type, evidence role, gate membership, the meaning of an empty value, and observed population counts. An empty cell is not automatically a missing measurement — several fields are conditional on task family, and their empty-value meaning is recorded explicitly.

Verification

This corpus is derived from sanitized publication bundles produced by the quant_eval harness. It is designed to be checked rather than trusted:

  • source_bundle_checksums.json, included here, republishes, verbatim, the SHA-256 digest and byte length of every file in every source bundle. No source file was modified.
  • Before this file was written, the builder verified all 72 source-file digests and independently recomputed all 96 family x runner pass rates from the raw per-case rows, matching the harness rollups exactly.
  • Aggregates published elsewhere in this corpus — the paired degradation statistics (D5) and family pass rates (D6) datasets — are recomputable from the per-case results dataset (D1) using the gate definitions in data_dictionary.json, included here.

Limits you should know before using this

  • Decoding conditions are not uniform across models. Temperature follows each publisher's own model card, so cross-model comparison of absolute pass rates is confounded. Within-run pairing is unaffected, which is what the paired test requires. The conditions are published per row and per run so they can be filtered on.
  • Runs on the Modal substrate record `seed` status `unsupported` — the deployed method signature accepts no seed — so those runs are not exactly reproducible. Local runs applied a fixed seed.
  • Runtime figures are observed harness wall time on the recorded hardware and backends. They are not a controlled throughput benchmark and not a general claim about quantization performance at any precision on any hardware. Direction is published as an explicit label because not every measured pair is a speedup.
  • The fuzz family is an adaptive trajectory evaluated from identical starting fixtures. Its paired test compares complete case outcomes, not identical post-divergence prompts.
  • Calibration runs are not published. Runs that informed a published run are disclosed by identifier in calibration_lineage.csv, published in the run provenance dataset (D4), so the record is complete without releasing provisional numbers.

Citation

bibtex
@dataset{pbh_quant_eval_d2,
  author    = {Hill, Patrick},
  title     = {quant_eval Throughput telemetry},
  publisher = {PBH Applied Systems, LLC},
  year      = {2026},
  doi       = {10.5281/zenodo.22009987},
  note      = {Version DOI: 10.5281/zenodo.22009988},
  license   = {CC-BY-4.0}
}

Licence

Creative Commons Attribution 4.0 International (CC BY 4.0). See LICENSE. Commercial use is permitted; attribution is required.

This corpus describes third-party models and redistributes no model weights. Each evaluated model remains under its own licence, recorded per run in the run provenance dataset.


Produced by builddatasets.py 2.4.0 from quanteval publication bundles. Built 2026-08-19.