Skip to content
#

gb10

Here are 176 public repositories matching this topic...

Serve DeepSeek and Qwen, run an on-demand model library, and fine-tune with Unsloth QLoRA on 2x NVIDIA DGX Spark. TP=2 vLLM lanes behind one OpenAI-compatible endpoint for OpenCode, Cursor, and Hermes. Honest benchmarks, committed artifacts.

  • Updated Aug 31, 2026
  • Python

Headless 4K remote desktop for the NVIDIA DGX Spark (GB10): one-command installer for Sunshine + Moonlight low-latency game streaming with NVENC hardware encoding, a software virtual display (no HDMI dummy plug), GDM autologin, and optional Tailscale.

  • Updated Jun 3, 2026
  • Shell

Serving Qwen3.8-27B-FP8 on a single DGX Spark (GB10): 7.88 to 58.5 tok/s single-stream from decode strategy alone, weights untouched. Speculative decoding and prefix caching benchmarked, plus DFlash 2 — the only Qwen3.8-27B build that can serve it under vLLM.

  • Updated Aug 19, 2026
  • Python

Serving 4-bit Qwen3.8-27B on a single DGX Spark (GB10): 75 tok/s single-stream, 246 tok/s aggregate at 8-way concurrency. NVFP4 vs MixedInt4-AutoRound vs the FP8 baseline, measured on one harness — including why the quantization advantage collapses to +0.2% by c16.

  • Updated Aug 19, 2026
  • Python

Improve this page

Add a description, image, and links to the gb10 topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the gb10 topic, visit your repo's landing page and select "manage topics."

Learn more