AMD FastFlowLM FastFlowLM
Models Benchmarks
Docs
Models Benchmarks
Docs
GitHub Discord YouTube Email

Docs

Instructions

Overview
Install-Windows
Install-Linux

Instructions

Overview Sys Command and CLI Mode Server Mode Server Basics API / Client Usage Open WebUI Tool Calling LangChain RAG LangChain Web Search Obsidian Microsoft AI Toolkit

Models

Overview LLaMA DeepSeek Qwen Gemma MedGemma TranslateGemma gpt-oss LiquidAI/LFM Microsoft/Phi Nanbeige Whisper EmbeddingGemma

Benchmarks

Overview LLaMA3 Gemma3 Gemma4 Qwen2.5 Qwen3 Qwen3.5 Qwen3.6 gpt-oss LiquidAI/LFM2 Microsoft/Phi4 Nanbeige4.1

🛠️ Instructions

FastFlowLM (FLM) is a deeply optimized runtime for local LLM inference on AMD NPUs —
ultra-fast, power-efficient, and 100% offline.

Its user interface and workflow are similar to Ollama, but purpose-built for AMD’s XDNA architecture.

This section will walk you through how to use FastFlowLM with examples.


📚 Sections

  • System Command and CLI Mode
  • Server Mode
  • Server Basics
  • API / Client Usage
  • Open WebUI
  • Tool Calling
  • LangChain RAG
  • LangChain Web Search
  • Obsidian
  • Microsoft AI Toolkit
AMD

About AMD

AMD (NASDAQ: AMD) drives innovation in high-performance and AI computing to solve the world's most important challenges. Today, AMD technology powers billions of experiences across cloud and AI infrastructure, embedded systems, AI PCs and gaming. With a broad portfolio of AI-optimized CPUs, GPUs, networking and software, AMD delivers full-stack AI solutions that provide the performance and scalability needed for a new era of intelligent computing. Learn more at www.amd.com.

Video Channels

AMD Developers Gaming

Discord

Developers Gaming

©2026 Advanced Micro Devices, Inc.

Terms and Conditions Privacy Trademarks Supply Chain Transparency Fair & Open Competition UK Tax Strategy Cookies Policy Cookie Settings