CoolFace
Modelpublic

devoppro/FastLLM

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes353downloads
Model Card

FastLLM (150M) — Modern Causal Language Model

FastLLM is a ~150M parameter, decoder-only causal language model built completely from scratch in PyTorch and fully integrated with Hugging Face transformers. It incorporates state-of-the-art LLM architectural choices—Grouped-Query Attention (GQA), SwiGLU MLPs, RMSNorm, and Rotary Position Embeddings (RoPE)—and natively saves weights in the zero-copy Safetensors format.


Model Details

  • —Developed by: devoppro
  • —Model Type: Decoder-only Causal Language Model
  • —Architecture: Custom Transformer (ModernLLMForCausalLM)
  • —Parameter Count: ~150,000,000 (150M)
  • —Tokenizer: Qwen 2.5 BPE Vocabulary (vocab_size: 151,936)
  • —Precision: Mixed Precision (FP16)
  • —Storage Format: .safetensors
  • —Repository: devoppro/FastLLM

Architectural Specifications

ParameterConfiguration
Hidden Size ($d_{\text{model}}$)768
Intermediate Size (SwiGLU)2048
Hidden Layers12
Query Heads12
Key/Value Heads (GQA)4 (3:1 Query-to-KV ratio)
Max Context Length2048 tokens
NormalizationRMSNorm ($\epsilon = 10^{-6}$)
Positional EmbeddingRotary Embeddings (RoPE, $\theta = 1000000.0$)

Training Data Mixture

The model was pre-trained using dynamic stream interleaving across four high-quality datasets: