ubr-physical-ai/Cosmos3-Edge-INT4-AWQ
Add measured accuracy: INT4 indistinguishable from bf16 once k is re-fitted
Warn that thinking is off by default and breaks structured output (11/24 vs 24/24 parseable)
Warn that thinking is off by default and breaks structured output (11/24 vs 24/24 parseable)
Card fixes: seven staging gaps not five; consistent GiB/GB units for the checkpoint
Card fixes: seven staging gaps not five; consistent GiB/GB units for the checkpoint
First confirmed Jetson Orin Nano bring-up: 52.8 tok/s, 2.24 GiB engines, 2.7 GB peak
First confirmed Jetson Orin Nano bring-up: 52.8 tok/s, 2.24 GiB engines, 2.7 GB peak
Correct the reasoner subset: und_prefill is a policy component, not part of the VLM path
Correct the reasoner subset: und_prefill is a policy component, not part of the VLM path
Fix the decoder MLP: 56 all-zero tensors were never bound (mlp.fc1/fc2 -> up_proj/down_proj)
Publish the v2 checkpoint: projector excluded from quantisation
Publish the v2 checkpoint: projector excluded from quantisation
Publish the v2 checkpoint: projector excluded from quantisation
Replace the broken ONNX export: real INT4 decoder (1.4 GB, 169 INT8 tensors), projector excluded
Correct the fit table and staging gaps: the 8.4 GB figure came from a broken export
Generalise Jetson notes to Orin Nano 4 GB / 8 GB
Generalise Jetson notes to Orin Nano 4 GB / 8 GB
Generalise Jetson notes to Orin Nano 4 GB / 8 GB
Add the Edge-LLM ONNX export (llm, visual, und_prefill, text_tokenizer); state the loading runtime, correct the 8-bit tag, add intended use and limitations
INT4 AWQ quantisation of nvidia/Cosmos3-Edge for Jetson Orin (unvalidated)
initial commit
