prathamkode/particle-1.0
2236
particle-1.0
~100M-parameter Llama-style chat model trained from scratch (random init). Not a fine-tune of Llama, SmolLM, or any Hub base.
Weights are MIT. Training data still needs attribution (below).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "prathamkode/particle-1.0"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
messages = [{"role": "user", "content": "hello"}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=64, temperature=0.7)
print(tok.decode(out[0], skip_special_tokens=False))Chat format:
<|user|>
hello
<|assistant|>Model details
Training
- Tokenizer trained from scratch on a FineWeb-Edu sample (~2GB text).
- Pretrain next-token prediction on `HuggingFaceFW/fineweb_edu_100BT-shuffled`, first ~2B tokens.
- SFT on `HuggingFaceTB/smol-smoltalk` (first user/assistant turn + a few greeting seeds).
SFT used that dataset as text only. No teacher model weights were copied.
Intended use
Research / demo small chat model. Expect short replies, mistakes, and weak reasoning.
Limitations
- Very small capacity
- May hallucinate
- English-centric FineWeb-Edu subset
- No RLHF / preference tuning
License
- These weights: MIT
- FineWeb-Edu: ODC-By (attribute)
- smol-smoltalk: follow the dataset card
