nineninesix/qwen3_5-full-attn-only-14
01k
Rebuilt Qwen3.5 — qwen3_5-full-attn-only
Note: This model was assembled with qwen3_5-rebuild. Embedding and output projection weights are copied from the source model. Transformer blocks are freshly initialized (random, std=0.02). The model requires fine-tuning before it can generate coherent text.Source model
KaniTTS-research-team/qwen-3.5-prepare-0.6b
Architecture
Parameter budget
embed_tokens represents 49.7% of total parameters. It is the dominant component — 254M params encode 248,320 token embeddings × 1024 dims.Layer-by-layer breakdown
Layer type sequence
0: full_attention
1: full_attention
2: full_attention
3: full_attention
4: full_attention
5: full_attention
6: full_attention
7: full_attention
8: full_attention
9: full_attention
10: full_attention
11: full_attention
12: full_attention
13: full_attentionInitialization
- Embeddings: copied as-is from
KaniTTS-research-team/qwen-3.5-prepare-0.6b - Transformer blocks:
randominit, std=0.02 - dtype:
bfloat16 - Seed:
42
