Pure Rust Inference Engine
-
Updated
Aug 31, 2026 - Rust
Pure Rust Inference Engine
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker issues, and validates vLLM/Ollama/llama.cpp/SGLang recipes.
Serve DeepSeek and Qwen, run an on-demand model library, and fine-tune with Unsloth QLoRA on 2x NVIDIA DGX Spark. TP=2 vLLM lanes behind one OpenAI-compatible endpoint for OpenCode, Cursor, and Hermes. Honest benchmarks, committed artifacts.
DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
Headless 4K remote desktop for the NVIDIA DGX Spark (GB10): one-command installer for Sunshine + Moonlight low-latency game streaming with NVENC hardware encoding, a software virtual display (no HDMI dummy plug), GDM autologin, and optional Tailscale.
Serving Qwen3.8-27B-FP8 on a single DGX Spark (GB10): 7.88 to 58.5 tok/s single-stream from decode strategy alone, weights untouched. Speculative decoding and prefix caching benchmarked, plus DFlash 2 — the only Qwen3.8-27B build that can serve it under vLLM.
Serving 4-bit Qwen3.8-27B on a single DGX Spark (GB10): 75 tok/s single-stream, 246 tok/s aggregate at 8-way concurrency. NVFP4 vs MixedInt4-AutoRound vs the FP8 baseline, measured on one harness — including why the quantization advantage collapses to +0.2% by c16.
Measured SM121 compatibility recipe for MiniMax H3 FL2VA on one NVIDIA DGX Spark with vLLM-Omni and online FP8.
Browser-based model and inference management for NVIDIA DGX Spark - inventory local and Hugging Face models, manage Ollama and LiteLLM, generate Docker Compose deployments for vLLM, SGLang, llama.cpp, LocalAI, and ComfyUI, with multi-user access, diagnostics, and multi-node support.
llama.cpp fork optimized for NVIDIA DGX Spark / GB10 (Blackwell, SM 12.1) — TurboQuant weights + KV, NVFP4, DFlash MTP
Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 network layer
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
Serve Poolside Laguna S 2.1 (NVFP4) on the NVIDIA DGX Spark (GB10) without hanging your box. Working stack, crash-safe configs, benchmarks, and the exact gotchas.
Add a description, image, and links to the gb10 topic page so that developers can more easily learn about it.
To associate your repository with the gb10 topic, visit your repo's landing page and select "manage topics."