RichardErkhov/abacaj_-_llama-161M-100B-gguf
0165
Quantization made by Richard Erkhov.
llama-161M-100B - GGUF
- Model creator: https://huggingface.co/abacaj/
- Original model: https://huggingface.co/abacaj/llama-161M-100B/
Original model description: --- library_name: transformers license: apache-2.0 ---
llama-161M
Trained on 100B tokens.
- 1e-3 LR
- 0.1 wd
- WSD scheduler with 10% decay
- 80% code, 10% NL, 10% instruction data
- Dataset decontaminated against popular benchmarks following bigcode
- 8x3090s 110~ hours
This is a base pretrained model and requires further fine tuning to be useful.
