工厂智能体如何通过子网络选择在有限硬件上高效运行,实测功耗仅1.3-5瓦。
研究显示,工厂现场助手模型经过结构压缩和检索增强适应后,模型大小不再是适应后答案质量的可靠预测指标。在制造业手册案例研究中,提取成本仅为未修剪模型判断质量的13.7%,检索增强蒸馏将其恢复到4.6%以内,恢复了三分之二的损失。同一助手可在三个异构边缘层级运行,待机功耗为1.3至5瓦。
Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents
On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no longer a reliable predictor of adapted answer quality: general capability falls almost linearly with parameter count, while judged retrieval-augmented answer quality does not. We therefore treat deployment as a post-adaptation selection problem, committing one sub-network per device on judged answer quality and measured on-device throughput under a configurable general-capability floor and memory budget; rules that optimize size, speed, or quality alone each give up capability or throughput. A weight-shared supernetwork trained with sandwich-style in-place distillation keeps this selection inexpensive. In a manufacturing-manual case study, extraction costs 13.7 percent of the unpruned model's judged quality and retrieval-grounded distillation returns it to within 4.6 percent, recovering two thirds of the loss, and the same assistant runs across three heterogeneous edge tiers at 1.3 to 5 watts standby.