⚡ Performance and Efficiency Benchmarks
This section reports the performance of Qwen 3.8 on NPU with FastFlowLM (FLM).
Note:
- Results are based on FastFlowLM v1.0.7.
- Under FLM’s default NPU power mode (Performance)
- Newer versions may deliver improved performance.
- Fine-tuned models show performance comparable to their base models.
qwen3.8-mtp:27bdecodes with its built-in MTP draft head, so decoding speed is prompt-dependent — predictable text (code, structured output, long reasoning chains) accepts more drafts and decodes faster than highly creative text. See the model card for how speculation works.- Benchmarks for this model are capped at 8k context.
Test System 1:
AMD Ryzen™ AI 5 340 (Kraken Point) with 64 GB DRAM; performance is comparable to other Kraken Point systems.
🚀 Decoding Speed (TPS, or Tokens per Second, starting @ different context lengths)
| Model | HW | 1k | 2k | 4k | 8k |
|---|---|---|---|---|---|
| Qwen3.8-27B | NPU (FLM) | 1.06 | 1.86 | 1.50 | 1.32 |
🚀 Prefill Speed (TPS, or Tokens per Second, with different prompt lengths)
| Model | HW | 1k | 2k | 4k | 8k |
|---|---|---|---|---|---|
| Qwen3.8-27B | NPU (FLM) | 104.94 | 110.56 | 113.05 | 112.81 |
🚀 Time to First Token (TTFT, Seconds, with different prompt lengths)
| Model | HW | 1k | 2k | 4k | 8k |
|---|---|---|---|---|---|
| Qwen3.8-27B | NPU (FLM) | 9.35 | 17.62 | 34.34 | 68.71 |
🚀 Prefill TTFT with Image Input (Seconds)
Prefill time-to-first-token (TTFT) for Qwen3.8-27B on NPU (FastFlowLM) with different image resolutions.
Mid Resolution Images:
| Model | HW | 720p (1280×720) | 1080p (1920×1080) |
|---|---|---|---|
| Qwen3.8-27B | NPU (FLM) | 15.4 | 34.7 |
This test uses a short prompt: “Describe this image.”