⚡ Performance and Efficiency Benchmarks

This section reports the performance of Qwen 3.6 on NPU with FastFlowLM (FLM).

Note:

  • Results are based on FastFlowLM v0.9.45.
  • Under FLM’s default NPU power mode (Performance)
  • Newer versions may deliver improved performance.
  • Fine-tuned models show performance comparable to their base models.

Test System 1:

AMD Ryzen™ AI 7 350 (Kraken Point) with 96 GB DRAM; performance is comparable to other Kraken Point systems.


🚀 Decoding Speed (TPS, or Tokens per Second, starting @ different context lengths)

Model HW 1k 2k 4k 8k 16k 32k
Qwen3.6-35B-A3B NPU (FLM) 13.65 13.41 13.09 12.51 11.24 9.51

🚀 Prefill Speed (TPS, or Tokens per Second, with different prompt lengths)

Model HW 1k 2k 4k 8k 16k 32k
Qwen3.6-35B-A3B NPU (FLM) 78.98 118.04 156.43 197.93 218.84 221.96