CoolFace
Modelpublic

nota-ai/st-llama-1-5.5b-ppl

sourceHugging Faceupdated 2y agoView on Hugging Face
10likes33downloads
Model Card

Shortened LLaMA Model Card

Shortened LLaMA is a depth-pruned version of LLaMA models & variants for efficient text generation.

  • —Developed by: Nota AI
  • —License: Non-commercial license
  • —Repository: https://github.com/Nota-NetsPresso/shortened-llm
  • —Paper: https://arxiv.org/abs/2402.02834

Compression Method

After identifying unimportant Transformer blocks, we perform one-shot pruning and light LoRA-based retraining. <details> <summary> Click to see a method figure. </summary>

<img alt="method" img src="https://netspresso-research-code-release.s3.us-east-2.amazonaws.com/compressed-llm/st-llama_method.png" width="100%">

</details>

Model Links

Source<br>ModelPruning<br>RatioPruning<br>CriterionHF Models<br>Link
LLaMA-1-7B20%PPLnota-ai/st-llama-1-5.5b-ppl
LLaMA-1-7B20%Taylor+nota-ai/st-llama-1-5.5b-taylor
Vicuna-v1.3-7B20%PPLnota-ai/st-vicuna-v1.3-5.5b-ppl
Vicuna-v1.3-7B20%Taylor+nota-ai/st-vicuna-v1.3-5.5b-taylor
Vicuna-v1.3-13B21%PPLnota-ai/st-vicuna-v1.3-10.5b-ppl
Vicuna-v1.3-13B21%Taylor+nota-ai/st-vicuna-v1.3-10.5b-taylor

Zero-shot Performance & Efficiency Results

  • —EleutherAI/lm-evaluation-harness version 3326c54

<img alt="results" img src="https://netspresso-research-code-release.s3.us-east-2.amazonaws.com/compressed-llm/st-llamazero-shotscores.png" width="100%">

License

  • —All rights related to this repository and the compressed models are reserved by Nota Inc.
  • —The intended use is strictly limited to research and non-commercial projects.

Acknowledgments

Citation

bibtex
@article{kim2024shortened,
  title={Shortened LLaMA: A Simple Depth Pruning for Large Language Models},
  author={Kim, Bo-Kyeong and Kim, Geonmin and Kim, Tae-Ho and Castells, Thibault and Choi, Shinkook and Shin, Junho and Song, Hyoung-Kyu},
  journal={arXiv preprint arXiv:2402.02834},      
  year={2024},
  url={https://arxiv.org/abs/2402.02834}
}
bibtex
@article{kim2024mefomo,
  title={Shortened LLaMA: A Simple Depth Pruning for Large Language Models},
  author={Kim, Bo-Kyeong and Kim, Geonmin and Kim, Tae-Ho and Castells, Thibault and Choi, Shinkook and Shin, Junho and Song, Hyoung-Kyu},
  journal={ICLR Workshop on Mathematical and Empirical Understanding of Foundation Models (ME-FoMo)},
  year={2024},
  url={https://openreview.net/forum?id=18VGxuOdpu}
}