BEE-spoke-data/smol_llama-220M-GQA
13510
smol_llama: 220M GQA
A small 220M param (total) decoder model. This is the first version of the model.
- 1024 hidden size, 10 layers
- GQA (32 heads, 8 key-value), context length 2048
- train-from-scratch on one GPU :)
Links
Here are some fine-tunes we did, but there are many more possibilities out there!
- instruct
- openhermes - link
- open-instruct - link
- code
- python (pypi) - link
- zephyr DPO tune
- SFT - link
- full DPO - link
Open LLM Leaderboard Evaluation Results
Detailed results can be found here
Open LLM Leaderboard Evaluation Results
Detailed results can be found here
