Touchstone:用可检验性判定哪些网络自动化任务能交给本地小模型
Can You Check That? The Checkability Boundary for Local LLM Network Automation
做网络运维的朋友看看:1-8B 小模型本地跑,配合自动检查规则,98.6% 准确率还不用把生产配置发给云端。
arXiv 论文提出 checkability 标准:任务若存在廉价、确定性的内在检查,就适合本地推理。作者实现 Touchstone 管线,用 7 个 1-8B 参数的开源小模型生成候选,再按任务内在检查过滤,未通过的输入升级给前沿大模型。在冲突检测和意图翻译两个任务上分别达到 98.6% 和 93.8% 准确率,升级率只有 16% 和 17%。而在没有内在检查的 TeleQnA 知识问答对照任务上,Touchstone 无法匹敌前沿模型基线。
Can You Check That? The Checkability Boundary for Local LLM Network Automation
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while escalating only 16% and 17% of inputs, respectively. On TeleQnA, a knowledge-only control that has no task-specific intrinsic checks, Touchstone is unable to match the accuracy of the frontier baseline. Our results support a simple deployment rule: keep inference local when task semantics support precise, low-cost checks; escalate the rest.