tencent/WeDLM-8B-Base
18198
WeDLM-8B
WeDLM-8B is a diffusion language model that performs parallel decoding under standard causal attention, initialized from Qwen3-8B.
This is the base (pretrained) version. For the instruction-tuned version, see WeDLM-8B-Instruct.
๐ Paper (Coming Soon) | ๐ Project Page | ๐ป GitHub
Model Details
Quick Start (Recommended)
For fast inference, use the wedlm engine:
pip install git+https://github.com/tencent/WeDLM.gitfrom wedlm import LLM, SamplingParams
llm = LLM(model="tencent/WeDLM-8B")
prompt = "The theory of relativity states that"
outputs = llm.generate([prompt], SamplingParams(max_tokens=256))
print(outputs[0]["text"])HuggingFace Transformers
For training or simple forward passes, you can load via Transformers:
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("tencent/WeDLM-8B", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
"tencent/WeDLM-8B",
trust_remote_code=True,
torch_dtype="auto",
device_map="auto"
)
inputs = tokenizer("The theory of relativity", return_tensors="pt").to(model.device)
outputs = model(**inputs)โ ๏ธ Note: The HuggingFace interface is for training/forward pass convenience. For optimized inference throughput, use the wedlm engine above.Performance
Citation (Coming soon)
License
Apache 2.0
