Skip to content
#

flash-next

Here are 2 public repositories matching this topic...

The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

  • Updated Sep 26, 2026
  • Python

Qwen3.8-Flash-Next 177B MoE on RTX 5090 Laptop (24GB VRAM + 64GB RAM): llama.cpp 22-25 tok/s -> Strata 93.5 tok/s (3.7x), GPU+CPU both saturated. 7 quant tiers screened over 26 rounds, MTP corruption found & fixed, 256K context + vision. 中英双语实测实录

  • Updated Oct 1, 2026
  • Python

Add this topic to your repo

To associate your repository with the flash-next topic, visit your repo's landing page and select "manage topics."

Learn more