CoolFace
Modelpublic

Se00n00/TinyLM2-50M-IFT

sourceHugging Facemitupdated 18d agoView on Hugging Face
0likes231downloads
Model Card

TinyLM2-50M-Instruct

TinyLM2-50M-Instruct is a compact decoder-only Transformer language model designed for efficient instruction following and conversational AI. The model has approximately 50M parameters and has been fine-tuned using Supervised Fine-Tuning (SFT) on the UltraChat 200k instruction-response dataset to improve chat capabilities while maintaining a lightweight footprint suitable for local inference and resource-constrained environments.

Evaluation

All evaluations are zero-shot unless stated otherwise, and i used lmeval to run them <img src="losschart.png">

Model Architecture & Hyperparameters

TinyLM2-50M-Instruct is built on a custom ALiBi Decoder-Only Transformer architecture with pre-normalization and gated feedforward networks:

HyperparameterValueDescription
ArchitectureALiBi Decoder-Only TransformerAutoregressive Decoder-Only Transformer
Total Parameters~50.96M (53,430,272)Compact and ultra-fast for edge & local CPU/GPU
inference
`vocab_size`50,271Includes special chat tags (`<SYSTEM>, <USER>, <ASSISTANT>`)
`hidden_size` (`d_model`)512Model hidden dimension
`intermediate_size` (`ff_hidden_d`)819SwiGLU Gated Feedforward hidden dimension
`num_hidden_layers`12Number of Transformer block layers
`num_attention_heads`8Attention heads (Head dim = 64)
`max_position_embeddings`2,048Maximum context sequence length
NormalizationRMSNorm (eps=1e-8)Scale normalization for accelerated throughput
Activation FunctionSwiGLU (SiLU)Gated Feedforward activation
Positional EncodingALiBiAttention with Linear Biases
Tie Word EmbeddingsTrueTied input embedding and LM head projection

## Tokenizer & Chat Template

The model uses a custom Byte-Level BPE Tokenizer equipped with special tokens and a pre-configured Jinja2 chat_template for multi-turn conversations.

PropertyValue
Tokenizer TypeGPT2Tokenizer (Byte-Level BPE)
Vocabulary Size50,271
Special Tokens`<START> <END> <UNK>`
Chat Control Tokens`<SYSTEM> <USER> <ASSISTANT>`
Extra Special Tokens`<THINK> </THINK> <AVAILABLE_TOOLS> </AVAILABLE_TOOLS> <TOOL_CALLS> </TOOL_CALLS> <TOOL_RESULTS> </TOOL_RESULTS>`
Chat TemplateNative Jinja2 support via tokenizer.apply_chat_template()

## Training Configuration

ParameterValue
Pipeline ProcessSupervised Instruction Fine-Tuning (SFT / IFT)
DatasetHuggingFaceH4/ultrachat_200k (train_sft, ~207k examples)
Total Examples~207k (4 epochs)
Learning Rate6e-5
Learning Rate ScheduleCosine (warmup_ratio=0.01)
Micro-Batch Size2 per device
Gradient Accumulation16 steps
Effective Batch Size32 × 2,048 tokens
OptimizerAdamW (weight_decay=0.1)
Max Sequence Length2,048 tokens
Precisionfloat16
HardwareNVIDIA Tesla T4 x 2 GPU

## Inference

python
  # pip install torch transformers
                                                                               
  import torch                                                                                            
  from transformers import pipeline                                                                       
                                                                                                          
  pipe = pipeline(                                                                                        
    "text-generation",                                                                                  
    model="Se00n00/TinyLM2-50M-Instruct",
    trust_remote_code = True                                                                            
  )                                                                                                       
                                                                                                          
  messages = [                                                                                            
    {"role": "system", "content": "You are a helpful AI assistant."},                                   
    {"role": "user", "content": "Explain artificial intelligence in simple terms."}                     
  ]                                                                                                       
                                                                                                          
  prompt = pipe.tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)       
                                                                                                          
  result = pipe(                                                                                          
    prompt,                                                                                             
    max_new_tokens=120,                                                                                 
    do_sample=True,                                                                                     
    temperature=0.5,                                                                                    
    top_k=40,                                                                                           
    top_p=0.9                                                                                           
  )                                                                                                       
                                                                                                          
  print(result[0]['generated_text'])                                                                  

──────

Sample Outputs

Chat-templated generation (temperature 0.8, top-p 0.9):

System: You are a helpful AI assistant

User: Write a short poem about the ocean.

The ocean is alive with life, As it floats in the waves, A vast expanse of vastness, A sight that cannot be matched, A wonderland that awakens, A place where the sun shines, The sea, it sets. The ocean is a reminder of life, The endless ocean, a place of peace, A place where the sun shines, The ocean is a beacon of hope, A place where the ocean is alive …

System: You are a helpful AI assistant

User: Explain the concept of gravity to a child.

The concept of gravity to a child is an important aspect of the child's development. It is the foundation of their physical and mental health. When a child is in a state of gravity, they are not able to move or move objects, which can lead to mental health problems such as anxiety, depression, and depression. The concept of gravity to a child is a crucial aspect of their development. It helps them to develop their mental health and develop their ability to control their own gravity. One of the main benefits of gravity is that …