🧩 Model Card: SmolVLA

  • Type: Vision-Language-Action
  • Think: No
  • Tool Calling Support: No
  • Base Model: lerobot/smolvla_base
  • Quantization: bf16
  • Max Context Length: 48 tokens
  • Max Camera Input: 3 images

📝 Note:

  • SmolVLA is a robotics policy model that maps camera images and language instructions directly to robot actions — it does not run in FLM’s standard CLI or Server chat modes.
  • For detailed usage instructions, please refer to: