devoppro/FastLLM
0353
FastLLM (150M) — Modern Causal Language Model
FastLLM is a ~150M parameter, decoder-only causal language model built completely from scratch in PyTorch and fully integrated with Hugging Face transformers. It incorporates state-of-the-art LLM architectural choices—Grouped-Query Attention (GQA), SwiGLU MLPs, RMSNorm, and Rotary Position Embeddings (RoPE)—and natively saves weights in the zero-copy Safetensors format.
Model Details
- Developed by: devoppro
- Model Type: Decoder-only Causal Language Model
- Architecture: Custom Transformer (
ModernLLMForCausalLM) - Parameter Count: ~150,000,000 (150M)
- Tokenizer: Qwen 2.5 BPE Vocabulary (
vocab_size: 151,936) - Precision: Mixed Precision (
FP16) - Storage Format:
.safetensors - Repository:
devoppro/FastLLM
Architectural Specifications
Training Data Mixture
The model was pre-trained using dynamic stream interleaving across four high-quality datasets:
