Gittensor Model Hub

Fast, practical checkpoints for single-GPU Blackwell — RTX 5090 and RTX PRO 6000 — quantized, benchmarked, and served by our own engine, SparkInfer.

Qwen3.8-27B-NVFP4-RTX5090

The flagship: NVIDIA ModelOpt NVFP4 of Qwen3.8-27B, tuned for GeForce RTX 5090.

  • Full 262K context on 32 GB
  • 81.6 tok/s alone — 155.8 tok/s with the drafter
  • Accuracy parity with Unsloth NVFP4

Qwen3.8-27B-DSpark-NVFP4

Matched DSpark drafter, trained and quantized against this exact checkpoint.

  • 1.91× decode, outputs byte-identical
  • 1.41 GB — a quarter of the built-in MTP head
  • +14.2% acceptance on agentic tool calling (v2)

SparkInfer

Blackwell-native, zero-dependency inference engine.

  • Custom RTX 5090 runtime
  • 2–3× faster than llama.cpp in many workloads

More