⚡ FastFlowLM (FLM)

FLM is the only NPU-first runtime built for AMD Ryzen™ AI.

Run LLMs — now with Vision support — in minutes: no GPU required, over 10× more power-efficient, and with context lengths up to 256k tokens.

Think Ollama — but laser-optimized for NPUs.

From idle silicon to instant powerFastFlowLM makes Ryzen™ AI shine.


📚 Sections

🚀 Get Started

Quick 5‑minute setup guide for Windows.

🐧 Get Started

Quick 5‑minute setup guide for Linux.

🛠️ Instructions

Run FastFlowLM using the CLI mode or server mode.

🧩 Models

Supported models, quantization formats, and compatibility details.

📊 Benchmarks

Real-time performance.