CoolFace
Datasetpublic

festr2/kimi-k3-aa-lcr-official-mxfp4-vs-qsrt-k2

Kimi-K3 AA-LCR official MXFP4 versus QSRT K2 evidence This repository stores a checksum-verified evidence archive for a paired comparison of the official Kimi-K3 MXFP4 checkpoint and the Kimi-K3 QSRT K2 routed-expert checkpoint. Evaluation conditions Dataset: ArtificialAnalysis/AA-LCR Dataset revision: bdae010bbce259820c0e34c1d7cce210d966fb75 Questions: 100 Independently generated answers per question and checkpoint: 3 Speculative decoding: disabled Sampling:… See the full description on the dataset page: https://huggingface.co/datasets/festr2/kimi-k3-aa-lcr-official-mxfp4-vs-qsrt-k2.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes24downloads
Dataset Card

Kimi-K3 AA-LCR official MXFP4 versus QSRT K2 evidence

This repository stores a checksum-verified evidence archive for a paired comparison of the official Kimi-K3 MXFP4 checkpoint and the Kimi-K3 QSRT K2 routed-expert checkpoint.

Evaluation conditions

  • —Dataset: ArtificialAnalysis/AA-LCR
  • —Dataset revision: bdae010bbce259820c0e34c1d7cce210d966fb75
  • —Questions: 100
  • —Independently generated answers per question and checkpoint: 3
  • —Speculative decoding: disabled
  • —Sampling: temperature=1.0, top_p=0.95, no seed, no system message
  • —Reasoning effort: max

The score is equality-checker accuracy over 300 attempts. It is not an official Artificial Analysis leaderboard result.

Results

Equality checkerOfficial MXFP4QSRT K2QSRT K2 minus official
Frozen official Kimi-K3254/300 (84.67%)245/300 (81.67%)-3.00 percentage points
GPT-5.6 Sol, maximum reasoning249/300 (83.00%)237/300 (79.00%)-4.00 percentage points

The frozen Kimi-K3 judge produced 18 official-only correct labels and 9 QSRT-only correct labels; the exact two-sided McNemar p-value is 0.1221. The GPT-5.6 Sol control produced 19 official-only and 7 QSRT-only correct labels; the exact two-sided McNemar p-value is 0.0290.

Archive

kimi-k3-aa-lcr-official-mxfp4-vs-qsrt-k2-20260815.tar.gz contains:

  • —all 600 generation receipts and raw API responses;
  • —all frozen Kimi-K3 and GPT-5.6 Sol judgement receipts;
  • —the official-judge repeatability control;
  • —both paired statistical comparison receipts;
  • —serving and generation manifests;
  • —the generation, judging, comparison, and packaging utilities;
  • —SHA-256 checksums for every archived file.

Archive SHA-256:

text
1a5ebe1adfc1249af1e9ebc1b49693346203c07ba5defb4dcb0bc9b16eb70ecc

After extraction, verify the archive from its root directory:

bash
sha256sum --check checksums.sha256

Checkpoint weights, AA-LCR source documents, and raw Docker inspection output are not included. Immutable checkpoint and dataset revisions, container image digests, runtime arguments, relevant non-secret environment variables, and source revisions are recorded in the archived manifests.