DominikBucko / qwen38-flash-next-2x3090 Star 100 Code Issues Pull requests Discussions Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt; prefill and agent 128K profiles), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode (64 GB profile). moe quantization mtp multi-gpu dual-gpu rtx3090 int4 fp8 vllm local-llm llm-inference qwen speculative-decoding rtx4090 rtx-3090 cpu-offloading qwen3-8 qwen38 qwen3-8-flash-next 256k-context Updated Sep 30, 2026 Python
lifeidle / qwen3.8-flash-next-strata-5090-laptop-24gb Star 13 Code Issues Pull requests Qwen3.8-Flash-Next 177B MoE on RTX 5090 Laptop (24GB VRAM + 64GB RAM): llama.cpp 22-25 tok/s -> Strata 93.5 tok/s (3.7x), GPU+CPU both saturated. 7 quant tiers screened over 26 rounds, MTP corruption found & fixed, 256K context + vision. 中英双语实测实录 moe strata quantization mtp vision-language-model llama-cpp local-llm gguf speculative-decoding rtx5090 qwen3 qwen38 256k-context flash-next Updated Oct 1, 2026 Python
mochgolf / qsa-hisparse Star 2 Code Issues Pull requests QSA HiSparse for SGLang: 256K KV offload, CUDA Graph benchmarks, and patches tested on dual RTX 4090 48GB qsa cpu-offload tensor-parallelism sparse-attention long-context fp8 gpu-inference llm-inference qwen sglang qwen3 rtx-4090 cuda-graphs sm89 256k-context hisparse rtx-4090-48gb 4090-48gb dual-rtx-4090 kv-cache-offloading Updated Sep 28, 2026 Python