高通用NPU跑agentic搜索,GPU要12秒CPU要30秒,NPU瞬间搞定还不发烫,这才是真有用的本地AI。
高通的Alan Zhu在Vector Space Day SF展示了设备端AI的延迟优势。NPU在agentic search任务中瞬时响应,而GPU耗时12秒,CPU耗时30秒后因发热降频。NPU在整个会话中保持90 tokens/秒,温度仅37°C。CPU速度持续下降。NPU的预填充速度足以在数秒内处理数千本地文件并返回引用答案,无需网络。
On-device AI has three advantages over cloud. Cost, privacy, and latency. Alan Zhu from @Qualcomm d...
On-device AI has three advantages over cloud. Cost, privacy, and latency. Alan Zhu from @Qualcomm demonstrated why latency is the most underestimated advantage of on-device AI at Vector Space Day SF. Alan ran the same agentic search task on NPU, GPU, and CPU. NPU responded instantly. GPU took 12 seconds. CPU took 30, then throttled down as the device heated up. The NPU held 90 tokens/sec for the full session at 37°C. The CPU started lower and kept dropping. The reason: NPU prefill speed. Fast enough to process thousands of local files and return a cited answer in seconds, entirely on device, no network required. That's what makes local AI agents actually practical, not just theoretically possible. Full talk: youtube.com/watch?v=FlAmmV… 💬 1 🔄 0 ❤️ 0 👀 60 📊 1 ⚡