⚡ FastFlowLM (FLM)
FLM is the only NPU-first runtime built for AMD Ryzen™ AI.
Run LLMs — now with Vision support — in minutes: no GPU required, over 10× more power-efficient, and with context lengths up to 256k tokens.
Think Ollama — but laser-optimized for NPUs.
From idle silicon to instant power — FastFlowLM makes Ryzen™ AI shine.
📚 Sections
🚀 Get Started
Quick 5‑minute setup guide for Windows.
🐧 Get Started
Quick 5‑minute setup guide for Linux.
🛠️ Instructions
Run FastFlowLM using the CLI mode or server mode.
🧩 Models
Supported models, quantization formats, and compatibility details.
📊 Benchmarks
Real-time performance.