LisaMegaWatts/MonarchSLM
0
Cache Monarch matrices + causal mask for faster inference
Fix completion_tokens: count tokens not decoded characters
Upload README.md with huggingface_hub
Upload Project.toml with huggingface_hub
Upload Dockerfile with huggingface_hub
Upload server.jl with huggingface_hub
Upload checkpoint.jl with huggingface_hub
Upload model.jl with huggingface_hub
initial commit
