wuj4s/qwen38-27b-chat-demo
Qwen3.8 Model Playground
A free, static chat interface for orcarouter/Qwen3.8-27B-Uncensored-FP8. This Space does not host model weights or provide inference credits. Connect an existing OpenAI-compatible endpoint serving the requested checkpoint to generate replies.
Features: streaming replies, multi-turn conversation, thinking display, temperature and output limits, editable system prompt, example prompts, cancellation, copy answer, and mobile layout. The UI clearly indicates when no endpoint is connected.
Connect inference
Choose Set up and enter your endpoint's HTTPS base URL (usually ending in /v1), served model name, and endpoint API key if required. Use the alias your server assigns to the requested checkpoint. Requests are sent directly from the browser to /chat/completions; the endpoint must support CORS. API keys and conversation history remain in tab memory and are cleared on refresh. Do not enter a Hugging Face account token. Your inference provider's charges and limits apply.
The author documents vLLM serving for this checkpoint in the model card. Its hosted OrcaRouter example currently names qwen/qwen3.8-27b, but the live OrcaRouter catalog describes that alias as the base model. Verify hosted model identity before using that alias for exact-checkpoint experiments.
Deployment status
The owner's free account was not eligible to create a ZeroGPU Space: Hugging Face requires PRO, or an eligible account at least 30 days old. This static frontend can be hosted without GPU billing. The separately prepared Gradio/PyTorch app can replace it when ZeroGPU hosting becomes available and a read-only model credential is added as a Space secret.
No inference backend or API key is bundled. Live model generation cannot be verified until an endpoint is configured. The checkpoint is an unaligned research model; see its model card for license, intended use, and limitations.
