Skip to content
#

256k-context

Here are 3 public repositories matching this topic...

Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt; prefill and agent 128K profiles), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode (64 GB profile).

  • Updated Sep 30, 2026
  • Python

Qwen3.8-Flash-Next 177B MoE on RTX 5090 Laptop (24GB VRAM + 64GB RAM): llama.cpp 22-25 tok/s -> Strata 93.5 tok/s (3.7x), GPU+CPU both saturated. 7 quant tiers screened over 26 rounds, MTP corruption found & fixed, 256K context + vision. 中英双语实测实录

  • Updated Oct 1, 2026
  • Python

Add this topic to your repo

To associate your repository with the 256k-context topic, visit your repo's landing page and select "manage topics."

Learn more