tencent/Sequential-Hidden-Decoding-8B-n4
1180
Sequential-Hidden-Decoding-8B-n4
This is the n=4 variant of Sequential Hidden Decoding, a method that scales sequence length by n× with only additional Embedding parameters — same Transformer, more compute per token.
- Base model: Qwen3-8B-Base
- Scale: 4×
- Additional Embedding Params: 3.1B
- Training Tokens: 150B
- Dtype: bfloat16
Note: This is a base model (not instruction-tuned). It is intended for benchmarking, text completion, and as a foundation for downstream fine-tuning (SFT / RLHF). For conversational or instruction-following use cases, please fine-tune on your own data.
Key Idea
Prepare n independent Embedding matrices to encode the same token sequence n times, interleave the results, and feed the n×-length sequence into the same Transformer. Only the last embedding of each token computes the next-token loss, while the preceding embeddings serve as implicit reasoning steps in a continuous latent space.
Results
Serving (SGLang)
This model requires a patched version of SGLang for inference. See the project page for installation options (Docker image, forked repo, or manual patch).
python -m sglang.launch_server \
--model-path tencent/Sequential-Hidden-Decoding-8B-n4 \
--trust-remote-code \
--tp-size 1 \
--port 30000 --host 0.0.0.0 \
--chunked-prefill-size -1 \
--attention-backend fa3 \
--mem-fraction-static 0.82 \
--max-running-requests 32 \
--context-length 131072 \
--cuda-graph-max-bs 128 \
--cuda-graph-bs 1 2 4 8 16 32 64 128from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.completions.create(
model="tencent/Sequential-Hidden-Decoding-8B-n4",
prompt="The meaning of life is",
max_tokens=128,
temperature=0,
)
print(response.choices[0].text)All Models
Citation
@article{hidden_decoding_2026,
title = {Hidden Decoding: Scaling Sequence Length in Pretraining},
year = {2026},
url = {https://welm.weixin.qq.com/posts/hidden_decoding/}
}License
This model is released under the License Terms of Sequential-Hidden-Decoding.
