Qwen3.8-27B-NVFP4-RTX5090
The flagship: NVIDIA ModelOpt NVFP4 of Qwen3.8-27B, tuned for GeForce RTX 5090.
- Full 262K context on 32 GB
- 81.6 tok/s alone — 155.8 tok/s with the drafter
- Accuracy parity with Unsloth NVFP4
Fast, practical checkpoints for single-GPU Blackwell — RTX 5090 and RTX PRO 6000 — quantized, benchmarked, and served by our own engine, SparkInfer.
The flagship: NVIDIA ModelOpt NVFP4 of Qwen3.8-27B, tuned for GeForce RTX 5090.
Matched DSpark drafter, trained and quantized against this exact checkpoint.
Blackwell-native, zero-dependency inference engine.
No-MTP variant (−0.85 GB) DSpark drafter · BF16 source Spark-Hermes-3.8-27B · weights soon