CoolFace
Modelpublic

fleetml/Qwen3.8-27B-Uncensored-DSpark-RTX5090

sourceHugging Faceapache-2.0updated 26d agoView on Hugging Face
0likes7.7kdownloads
Model Card

Qwen3.8-27B Uncensored DSpark NVFP4 RTX 5090

Target-matched ModelOpt NVFP4 DSpark speculative drafter for fleetml/Qwen3.8-27B-Uncensored-NVFP4-RTX5090.

The drafter was initialized from RadixArk/Qwen3.8-27B-DSpark, trained in BF16 on hidden states captured from the exact OrcaRouter-derived NVFP4 target, and converted to ModelOpt NVFP4 for RTX 5090 serving.

Combined profile

FieldValue
Combined disk size22.24 GB
Combined GPU weight allocation20.50 GB
GPU memory remaining after initialization2.83 GB
Architecture context window262,144 tokens
Validated launch context122,880 tokens
Automatically allocated active pool87,798 tokens

Measured result

The current NVFP4 pair reached 217.73 tokens per second median decode across 18 successful requests on one RTX 5090.

MetricResult
Median decode217.73 tokens per second
Median time to first token0.113 seconds
P95 time to first token0.149 seconds
Generated tokens per request26 to 640
Total generated tokens3,633

The test used concurrency 1, temperature 0, two warmups, three repeats across six prompts, FP8 E4M3 KV cache, and a fixed request seed. The 16,000-token request setting was a ceiling. This result is not a sustained 16,000-token generation measurement.

Drafter details

FieldValue
ArchitectureDSparkDraftModel
Serving precisionModelOpt NVFP4
Training precisionBF16
Parameters1,857,358,337
Layers5
Block size7
Target layers5, 19, 33, 47, 61
Export size1,642,153,266 bytes
Safetensors SHA-256cf5720d722e0b0bbd74198262c8857ee3a348a14e99ab30bd96cb31e353b5ebe

Training and conversion

  • —Exact-target teacher features: 126 records
  • —Maximum training sequence length: 1,024 tokens
  • —Training length: 64 optimizer steps
  • —Dataset source: HuggingFaceH4/ultrachat_200k
  • —Training framework: SpecForge at commit 2fc993077c3c53df14ebdd5d414c854864476af8
  • —Quantization: NVIDIA ModelOpt NVFP4
  • —Conversion validation: tensor, dtype, metadata, load, API, Hermes, and benchmark checks

Serving

Use this repository as --speculative-draft-model-path and set --speculative-draft-model-quantization modelopt_fp4. The complete RTX 5090 launch command is provided on the target model page.

Use

This drafter is intended for controlled local research, evaluation, and agent development. Use it responsibly and comply with the Apache 2.0 license and applicable law.