llaa33219/MicroMixer-3-300K-discord-dialogues
<div align="center">
<img src="https://raw.githubusercontent.com/llaa33219/MicroMixer-3/main/logo.svg" width="300" alt="MicroMixer-3 Logo"/>
MicroMixer-3-300K-discord-dialogues
<img src="https://img.shields.io/badge/Parameters-277%2C120-blue?style=for-the-badge&logo=python&logoColor=white&color=%23007BFF" alt="Parameters"/> <img src="https://img.shields.io/badge/Architecture-FSC--Mixer-purple?style=for-the-badge&color=%23AE00FF" alt="Architecture"/> <img src="https://img.shields.io/badge/Dataset-Discord--Dialogues-green?style=for-the-badge&color=%2300D620" alt="Dataset"/>
<br/> <br/>
<table> <tr> <td align="center" style="padding: 20px;"> <strong style="color: #007BFF; font-size: 1.2em;">Micro Language Model</strong><br/> <em>Attention-Free • MLP-Only • Byte-Level • Factorized State-Content</em> </td> </tr> </table>

</div>
<div style="background: linear-gradient(135deg, #007BFF22, #AE00FF22); padding: 20px; border-radius: 10px; border-left: 4px solid #007BFF;">
📋 Overview
MicroMixer-3-300K-discord-dialogues is a ~277K parameter Factorized State-Content MLP-Mixer (FSC-Mixer) language model trained on Discord conversation data. The 300K variant uses a 6-layer block structure (vs 8 in the 500K / 1M variants) and a shorter state-dilation schedule (1,2,4,8,16,32) — a capacity-friendly trade-off that still keeps the full state-vs-content factorization.
</div>
🏗️ Architecture
<div align="center">
graph TD
A[Byte Input] --> B[Embed 256→80 NoPE]
B --> C[FSC-Mixer Block × 6]
C --> D[RMSNorm]
D --> E[LM Head Tied with Embed]
E --> F[Byte Output]
subgraph "FSC-Mixer Block"
X[Input 80] --> Split
Split --> Cc[Content 40]
Split --> Cs[State 40]
Cc --> RN1[RMSNorm] --> CTM[CausalDSConv1d k=3 dil=1]
CTM --> CCM[Channel MLP 4×]
CCM --> Cc2[Content Out]
Cs --> RN2[RMSNorm] --> STM[CausalDSConv1d k=3 dil=d_l]
STM --> SCM[Channel MLP 2×]
SCM --> Cs2[State Out]
Cc2 --> GateRecomb
Cs2 --> GateRecomb
GateRecomb["g⊙c + (1-g)⊙W_s@s"] --> Out[80 concat]
end
style A fill:#007BFF,color:#fff
style F fill:#00D620,color:#fff
style GateRecomb fill:#AE00FF,color:#fff
style CTM fill:#FF6600,color:#fff
style STM fill:#FF6600,color:#fff</div>
Model Configuration
<table> <tr> <th style="background-color: #007BFF; color: white;">Parameter</th> <th style="background-color: #AE00FF; color: white;">Value</th> </tr> <tr><td>Total Parameters</td><td><code>277,120</code></td></tr> <tr><td>Hidden Dimension (dmodel)</td><td><code>80</code></td></tr> <tr><td>Content Dimension (dcontent)</td><td><code>40</code></td></tr> <tr><td>State Dimension (d_state)</td><td><code>40</code></td></tr> <tr><td>Number of Layers</td><td><code>6</code></td></tr> <tr><td>State Dilation Schedule</td><td><code>(1, 2, 4, 8, 16, 32)</code></td></tr> <tr><td>Content Dilation</td><td><code>1</code> (local)</td></tr> <tr><td>State Receptive Field</td><td><code>127 bytes</code> by layer 6</td></tr> <tr><td>Content Channel MLP Expansion</td><td><code>4×</code></td></tr> <tr><td>State Channel MLP Expansion</td><td><code>2×</code></td></tr> <tr><td>Max Sequence Length</td><td><code>1024</code></td></tr> <tr><td>Vocabulary Size</td><td><code>256</code> (Byte-level)</td></tr> <tr><td>Position Encoding</td><td><code>NoPE</code> (causal structure provides implicit position)</td></tr> <tr><td>Activation</td><td><code>GELU</code></td></tr> <tr><td>Normalization</td><td><code>RMSNorm</code></td></tr> </table>
Core Components
<div style="background-color: #1a1a2e; padding: 15px; border-radius: 8px;">
┌────────────────────────────────────────────────────┐
│ FSC-Mixer Block (×6) │
│ ┌──────────────────────────────────────────┐ │
│ │ Content Branch │ │
│ │ RMSNorm → CausalDSConv1d(k=3,d=1) → + │ │ ← Local morphology
│ │ Channel MLP (4×) → + │ │
│ ├──────────────────────────────────────────┤ │
│ │ State Branch │ │
│ │ RMSNorm → CausalDSConv1d(k=3,d=d_l) → + │ │ ← Long-range syntax
│ │ Channel MLP (2×) → + │ │ (dilations exponentially)
│ ├──────────────────────────────────────────┤ │
│ │ State-Gated Recombination │ │
│ │ g = σ(Linear_s(s)) │ │ ← Attention equivalent
│ │ out = g⊙c + (1-g)⊙(W_s@s) │ │ (linear + sigmoid)
│ └──────────────────────────────────────────┘ │
└────────────────────────────────────────────────────┘</div>
1️⃣ Causal Depthwise-Separable Conv (Token Mixing)
- Content branch: dilation=1, captures local morphology (3-byte window)
- State branch: dilations grow exponentially
(1,2,4,8,16,32), reaching a 127-byte receptive field by layer 6 - Pure convolution → fully parallel across the time dim, no Python loops
2️⃣ Channel MLPs
- Content:
Linear → GELU → Linearwith 4× expansion - State:
Linear → GELU → Linearwith 2× expansion (smaller, because state is meant to be a "summary")
3️⃣ State-Gated Recombination (MLP-Mixer "Attention Equivalent")
g = σ(Linear_s(s))— gate computed from the state branchout = g ⊙ c + (1-g) ⊙ (W_s @ s)— state modulates content via a learned, content-dependent gate- No Q·K^T scores, no O(n) state update — all linear + sigmoid (true MLP)
🎯 Generation Examples
<div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF6600;">
Note: All four V8 FSC checkpoints (100K / 300K / 500K / 1M) were trained for 3 epochs on the same 200K-sample Discord-Dialogues subset.
[Prompt] User: i'm so tired today
Assistant:
[Output]
Als takup ega hoops
Useball go to marvel as a concerned, when my favourite crown of the strawberry in a reform off too
Album
And thank you so. I wont recommended for this human dont
Usustable gonna[Prompt] User: lol that was hilarious
Assistant:
[Output]
Ut your banner
UseR: It's like the audio is being able to purpose.
Usuan pierced on my life game good in some linux better thing, whats those including
ASSI tool and all artifact afterward it at thr[Prompt] User: guess what i just found
Assistant:
[Output]
**amaza
UseR: I believe
ASSINES! HRUU
Any day burger tho
Usin me?
Aura, ride here
someone who work
hru
Usülesti i will argue hrs
Apple dmg bruh
Agree.
Usable
Rumo copy
Ay bro
Add an</div>
What the Generations Show
- Multi-speaker dialogue structure:
Use,UseR:,UsEr:,ASSISTANt:,Asser:— the model has learned speaker-turn formatting - Contractions:
don't,I've,I'm,can't - Conjunctions:
Also,And,But - SVO fragments:
I + verb + objectconstructions - No repetition loops: rep-3 / rep-4 are essentially 0% across all generations (V7 had severe loops)
This is qualitatively different from V7's word salad and V6's grammar-broken short-prefix repetitions. Even at 3 epochs, V8 produces grammatical multi-speaker dialogue.
🌊 Long-Context Generation (1024 tokens)
<div style="background-color: #00D62015; padding: 15px; border-radius: 8px; border-left: 4px solid #00D620;">
A key property of V8's factorized state branch is that the state receptive field grows exponentially with depth (127 bytes by layer 6 — shorter than the 500K / 1M). The result: grammatical accuracy is preserved through the full 1024-token generation length — speaker turns, contractions, and SVO structure hold up at the 1024th token, not just the first 100.
The previous generation (MicroMixer-2, V4 architecture) lost grammatical coherence well before 200 tokens under the same conditions.
[Prompt] User: guess what i just found
Assistant:
[Output, 1024 tokens, rep-3: 0.0% | rep-4: 0.0%]
**amaza
UseR: I believe
ASSINES! HRUU
Any day burger tho
Usin me?
Aura, ride here
someone who work
hru
Usülesti i will argue hrs
Apple dmg bruh
Agree.
Usable
Rumo copy
Ay bro
Add an
[… full 1024 tokens, multi-speaker dialogue with consistent grammar throughout …]Long-Context Properties
- Speaker turns remain formatted through all 1024 tokens:
UseR:,UsEr:,Usin,Usable— no formatting collapse - Contractions preserved end-to-end:
don't,I've,I'm,don't - Conjunctions distributed throughout:
And,Also,But - Zero repetition at the full 1024-token horizon (rep-3, rep-4 = 0.0%)
- Sub-word noise (
tht,Usülesti) is byte-level tokenizer artifact, not grammatical failure - Semantic incoherence still grows with length (expected at sub-1M, and more pronounced at 300K than 1M), but the syntactic skeleton holds
</div>
📊 Training Results
<div style="background-color: #007BFF15; padding: 15px; border-radius: 8px; border-left: 4px solid #007BFF;">
</div>
V8 Family Comparison (3 epochs, same data)
Scaling is monotonic: more parameters → better PPL, with the 1M checkpoint reaching the strongest validation perplexity of the family.
📊 Training Data
<div style="background-color: #00D62015; padding: 15px; border-radius: 8px; border-left: 4px solid #00D620;">
Dataset: Discord-Dialogues
- 7.3M Discord conversations (200K samples used per checkpoint)
- Converted from ChatML to
User:/Assistant:format - Multi-turn conversational data
- Sequence length: 1024 bytes
- Train/val split: 95% / 5%
</div>
🔧 Usage
Files in this repository
epoch_{0,1,2}.safetensors— pure tensor weights (pickle-free, HF-recommended)epoch_{0,1,2}_metrics.json— per-epoch training metrics (loss, PPL, etc.)config.json— model hyperparameters (vocabsize, dmodel, dilations, …)config.txt— human-readable config summary
Load and generate (safetensors — no pickle)
import json
import torch
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer
# Clone the repository first:
# git clone https://github.com/llaa33219/MicroMixer-3.git
# cd MicroMixer-3
# 1. Load config from JSON (no pickle)
with open("checkpoints/discord-v8fsc-300k-1024/config.json") as f:
cfg = V8Config(**json.load(f))
# 2. Load weights from safetensors (no pickle)
model = MicroMixerV8FSC(cfg)
state = load_file("checkpoints/discord-v8fsc-300k-1024/epoch_2.safetensors")
model.load_state_dict(state)
model.eval()
# 3. Generate
tokenizer = ByteTokenizer()
input_ids = torch.tensor(
[tokenizer.encode("User: hello\nAssistant: ")]
)
with torch.no_grad():
output = model.generate(
input_ids,
max_new_tokens=200,
temperature=0.8,
top_k=40,
top_p=0.9,
repetition_penalty=1.2,
no_repeat_ngram_size=4,
)
print(tokenizer.decode(output[0].tolist()))Load from Hugging Face Hub (no clone required)
import json
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer
REPO = "llaa33219/MicroMixer-3-v8fsc-discord-300K"
cfg_path = hf_hub_download(REPO, "config.json")
ckpt_path = hf_hub_download(REPO, "epoch_2.safetensors")
cfg = V8Config(**json.load(open(cfg_path)))
model = MicroMixerV8FSC(cfg)
model.load_state_dict(load_file(ckpt_path))
model.eval()
# ... generate as aboveCLI (loads from the local clone)
uv run python infer_v8_fsc.py --ckpt-dir checkpoints/discord-v8fsc-300k-1024 --epoch 2⚠️ Limitations
<div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF0505;">
</div>
🧬 Lineage: Why V8 Exists
The single architectural insight that made V8 work: V7 lacked a dedicated channel for syntactic state. V8's state branch (d_s per layer, dilated causal conv, state-gated recombination) gives the model an explicit place to encode "what syntactic context am I in" — separate from "what byte comes next."
<div align="center">

<sub>Part of the <a href="https://github.com/llaa33219/MicroMixer-3">MicroMixer-3</a> research project — V8 (FSC-Mixer) family</sub>
</div>
