Baekpica/K2-Horizon-375B-A23B-Mixed-Quant-GGUF
Round 6: pipelined worklist MMQ K loop (8e886f1) and ldmatrix HMMA attention (12a2e14): 644.78 / 641.94 tok/s prefill, byte-identical
Record round 5 (guarded routed-down output, 1024-token prefill chunks) DGX Spark results
Record round 4 (IQ1 gate/up pair, cached RoPE angles) DGX Spark results
Record round 3 (IQ2_XS worklist, native per-chunk decode attention) DGX Spark results
Record Decode 2 (split-K decode attention) DGX Spark results
Record Prefill 6 (IQ1_M MMQ tile) DGX Spark results
Update README.md
Record Decode 1 GQA tile-2 Spark numbers
Record Prefill 4/5 rejects on Spark campaign table
Record Prefill 3 IQ1_M slot-loop Spark numbers
Note Prefill 2 IQ2_XS worklist reject vs locked baseline
Record Prefill 1 IQ1_M assign-major Spark numbers
Add ds4 K2 H200 validation receipt
Document ds4 H200 validation and Spark follow-up
Upload validation/MQ87.remote.audit.json with huggingface_hub
Upload README.md with huggingface_hub
Upload MQ87-SHA256SUMS with huggingface_hub
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Add files using upload-large-folder tool
Publish MQ87 dry-run and imatrix consistency audits
Add H200 MQ87 runtime smoke gate
Add remote LFS checksum verifier
Record MQ87 dry-run gate and production status
Add fail-closed MQ87 split and audit finalizer
Add files using upload-large-folder tool
Document K2 tokenizer calibration path
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
initial commit
