Mieaz/gpt22m-chat
028
beautiful transparent model card: full training trace, benchmark, limitations, reproduction
add with-past ONNX for browser inference (24L GQA-4)
GPT-22M: 24L/256 GQA-4, vocab16k, stories->chat anneal
initial commit
beautiful transparent model card: full training trace, benchmark, limitations, reproduction
add with-past ONNX for browser inference (24L GQA-4)
GPT-22M: 24L/256 GQA-4, vocab16k, stories->chat anneal
initial commit