⚡ Performance and Efficiency Benchmarks
This section reports the performance of Qwen 3.6 on NPU with FastFlowLM (FLM).
Note:
- Results are based on FastFlowLM v0.9.45.
- Under FLM’s default NPU power mode (Performance)
- Newer versions may deliver improved performance.
- Fine-tuned models show performance comparable to their base models.
Test System 1:
AMD Ryzen™ AI 7 350 (Kraken Point) with 96 GB DRAM; performance is comparable to other Kraken Point systems.
🚀 Decoding Speed (TPS, or Tokens per Second, starting @ different context lengths)
| Model | HW | 1k | 2k | 4k | 8k | 16k | 32k |
|---|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | NPU (FLM) | 13.65 | 13.41 | 13.09 | 12.51 | 11.24 | 9.51 |
🚀 Prefill Speed (TPS, or Tokens per Second, with different prompt lengths)
| Model | HW | 1k | 2k | 4k | 8k | 16k | 32k |
|---|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | NPU (FLM) | 78.98 | 118.04 | 156.43 | 197.93 | 218.84 | 221.96 |