CoolFace
Modelpublic

afrideva/TinyLlama-1.1B-intermediate-step-1431k-3T-GGUF

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
1likes760downloads
Model Card

TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T-GGUF

Quantized GGUF model files for TinyLlama-1.1B-intermediate-step-1431k-3T from TinyLlama

Original Model Card:

<div align="center">

TinyLlama-1.1B

</div>

https://github.com/jzhang38/TinyLlama

The TinyLlama project aims to pretrain a 1.1B Llama model on 3 trillion tokens. With some proper optimization, we can achieve this within a span of "just" 90 days using 16 A100-40G GPUs ๐Ÿš€๐Ÿš€. The training has started on 2023-09-01.

<div align="center"> <img src="./TinyLlama_logo.png" width="300"/> </div>

We adopted exactly the same architecture and tokenizer as Llama 2. This means TinyLlama can be plugged and played in many open-source projects built upon Llama. Besides, TinyLlama is compact with only 1.1B parameters. This compactness allows it to cater to a multitude of applications demanding a restricted computation and memory footprint.

This Collection

This collection contains all checkpoints after the 1T fix. Branch name indicates the step and number of tokens seen.

Eval
ModelPretrain TokensHellaSwagObqaWinoGrandeARC_cARC_eboolqpiqaavg
Pythia-1.0B300B47.1631.4053.4327.0548.9960.8369.2148.30
TinyLlama-1.1B-intermediate-step-50K-104b103B43.5029.8053.2824.3244.9159.6667.3046.11
TinyLlama-1.1B-intermediate-step-240k-503b503B49.5631.4055.8026.5448.3256.9169.4248.28
TinyLlama-1.1B-intermediate-step-480k-1007B1007B52.5433.4055.9627.8252.3659.5469.9150.22
TinyLlama-1.1B-intermediate-step-715k-1.5T1.5T53.6835.2058.3329.1851.8959.0871.6551.29
TinyLlama-1.1B-intermediate-step-955k-2T2T54.6333.4056.8328.0754.6763.2170.6751.64
TinyLlama-1.1B-intermediate-step-1195k-token-2.5T2.5T58.9634.4058.7231.9156.7863.2173.0753.86