Liquid AI发布端侧视觉语言模型LFM2.5-VL-3B

Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device

精选理由

Liquid AI把视觉语言模型塞进端侧,能读屏、指物、调工具,3GB体积在苹果芯片上跑得飞快。

AI 摘要

Liquid AI发布LFM2.5-VL-3B,一个3.1B参数的视觉语言模型,专为端侧部署设计。该模型在ScreenSpot-v2上平均得分80.7,在RefCOCO指代分割任务上从57.1提升至87.9。函数调用能力首次加入VL系列,ToolSandbox得分从26.4跃升至59.5。模型体积约3GB,在Apple M5 Max上解码速度达228 tokens/s。

图片来源 · marktechpost
原文 · marktechpost

Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device

Liquid AI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model built for on-device deployment. It averages 80.7 on ScreenSpot-v2 and lifts RefCOCO grounding from 57.1 to 87.9. Function calling is new to the VL line, with ToolSandbox moving from 26.4 to 59.5. The model fits in roughly 3 GB and decodes 228 tokens/s on an Apple M5 Max. The post Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device appeared first on MarkTechPost .