fleetml/Qwen3.8-27B-Uncensored-DSpark-RTX5090
Qwen3.8-27B Uncensored DSpark NVFP4 RTX 5090
Target-matched ModelOpt NVFP4 DSpark speculative drafter for fleetml/Qwen3.8-27B-Uncensored-NVFP4-RTX5090.
The drafter was initialized from RadixArk/Qwen3.8-27B-DSpark, trained in BF16 on hidden states captured from the exact OrcaRouter-derived NVFP4 target, and converted to ModelOpt NVFP4 for RTX 5090 serving.
Combined profile
Measured result
The current NVFP4 pair reached 217.73 tokens per second median decode across 18 successful requests on one RTX 5090.
The test used concurrency 1, temperature 0, two warmups, three repeats across six prompts, FP8 E4M3 KV cache, and a fixed request seed. The 16,000-token request setting was a ceiling. This result is not a sustained 16,000-token generation measurement.
Drafter details
Training and conversion
- Exact-target teacher features: 126 records
- Maximum training sequence length: 1,024 tokens
- Training length: 64 optimizer steps
- Dataset source: HuggingFaceH4/ultrachat_200k
- Training framework: SpecForge at commit
2fc993077c3c53df14ebdd5d414c854864476af8 - Quantization: NVIDIA ModelOpt NVFP4
- Conversion validation: tensor, dtype, metadata, load, API, Hermes, and benchmark checks
Serving
Use this repository as --speculative-draft-model-path and set --speculative-draft-model-quantization modelopt_fp4. The complete RTX 5090 launch command is provided on the target model page.
Use
This drafter is intended for controlled local research, evaluation, and agent development. Use it responsibly and comply with the Apache 2.0 license and applicable law.
