mohith-das/jetson-hybrid-flat-cache-0.5b
044
Add proper model card
Add Q8_0 GGUF (GLA layers fall back to softmax attention in llama.cpp)
Phase 2 hybrid SWA/GLA model, CPT-stabilized, FP16
initial commit
Add proper model card
Add Q8_0 GGUF (GLA layers fall back to softmax attention in llama.cpp)
Phase 2 hybrid SWA/GLA model, CPT-stabilized, FP16
initial commit