0xSero/GLM-5.3-Flash-EXL3-TR3-3.0bpw
GLM-5.3-Flash EXL3 TR3 3.0 bpw
Campaign status: layerwise encoding pending
This public repository is reserved for an independent MIT-lineage TR3 adaptation of zai-org/GLM-5.3-Flash-BF16, pinned at a6c167b62691b2bac901344b65cb651a70f53e43.
No model weights are present yet. The repository will be promoted only after the real-layer oracle, full 42-layer/288-expert encode, structural assembly, CUDA-graph primitive test, five cold KLD runs, upload, and independent public manifest verification complete. Full-server generation, vision, and MTP remain separate gates and will not be inferred from a structural or offline quality pass.
Planned method and scope
Routed-expert gate/up/down projections in language layers 3-44 are quantized layerwise from BF16 to calibrated 3.0 bpw EXL3/Trellis with TP4 rank slicing, pooled scale search, and cross-slice lockstep LDLQ. Attention, routers, shared experts, dense layers 0-2, embeddings, LM head, norms, vision, and MTP stay at source precision.
The sealed 600 x 2,048-token calibration deterministically replaces 64 eligible natural-text rows with scrubbed samples derived from the operator's private sessions. All 92 protected random/diversity rows remain byte-exact. Raw private content is not included or published; only aggregate counts, verification receipts, and cryptographic digests will appear in the release.
Attribution and clean-room boundary
- Z.AI: MIT-licensed GLM-5.3-Flash BF16 base model.
- ExLlamaV3 / TurboDerp: EXL3 and Trellis implementation.
- Brandon M. Music: MIT-licensed GLM-5.2 TR3 v3.1 encoder and rank-sliced release lineage, pinned from brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw at commit
f79c9167690ca705e877ae4dc55a841d1aae1247. - Dione: independent GLM-5.3 BF16 capture, adaptation, validation, and release workflow.
No code, weights, calibration corpus, or generated artifact from Brandon Music's separately licensed GLM-5.3 release is used. This work is not endorsed by any credited party.
See RELEASE_STATUS.json and PROVENANCE.md for the current pending-state contract.
Release identity
Status: Placeholder; no weights, runtime or quality release. Audited weight payload: 0 safetensors files, 0 bytes (weight files only; excludes metadata).
Upstream source: zai-org/GLM-5.3-Flash-BF16, BF16 revision a6c167b62691b2bac901344b65cb651a70f53e43. Artifact/evidence snapshot inspected: `2720a0e3ace2f876f5872d8ad826ad40c3d3b572`. A card update does not constitute a new weight conversion.
Availability
No downloadable model is present in this repository. Calibration budgets and component layouts above describe the planned conversion; runtime, allocation and quality gates remain pending for this variant. The inventory below identifies populated alternatives.
Intended use and limitations
Use populated checkpoints for local inference or quantization research with the declared compatible runtime. Results from one bitrate or runtime do not transfer automatically to another. Quantization may change behavior and factual accuracy; controlled smoke tests do not establish broad benchmark quality. A projection or REAP observation record alone does not prove successful refusal removal or a pruned model release.
Related releases
Private links require authorized access. Two suite indexes and three placeholders are included in this inventory; they are not additional trained models.
REAP observation provenance
The original 3bpw and Q4 were each observed on two corpora: private calibration material and balanced 12-language Wikipedia. Each sealed lane records 128 sequences × 1,024 tokens (131,072 tokens), across 42 routed layers and 288 experts per layer. These four observation lanes are separate from quantization calibration and do not mean experts have been pruned from the weights above. The private observation dataset holds the manifests and aggregate sidecars.
