这篇论文全面梳理了LLM智能体在安全领域的应用现状,帮你了解技术进展和未来方向。
这篇论文系统综述了2023-2026年间基于大语言模型的智能体在软件和系统安全领域的应用。研究涵盖了智能体的技术架构、感知、记忆、推理和规划等七个方面。论文分析了不同安全任务中的应用场景,并评估了数据集、结果指标和安全措施。研究指出当前智能体已具备行动能力,但权限边界和行为审计仍存在挑战。
LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment
Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypotheses, invoke tools, interpret outputs, and revise plans. Large language model (LLM)-based agents, which can plan, use tools, retain state, and revise actions across multi-step workflows, are being rapidly adopted to automate this work. Given the consequences of delegating security decisions to autonomous systems, understanding how such agents are built, used, and assessed is crucial. Yet to this date, there remains a lack of systematic understanding of what has been done and how far we are in this field: the term "agent" is applied inconsistently, applications differ sharply in risk, and assessment protocols are often incomparable. To gain a comprehensive and coherent view of this area hence inform relevant future research, this paper provides a systematic literature review of the (1) technical approaches, including agent architecture, perception, memory, reasoning and planning, action space, orchestration, and self-improvement, (2) applications, with respect to the security tasks served, and (3) assessment, including the datasets, outcome and trajectory metrics, safety measures, and baselines considered, over the peer-reviewed literature spanning the emergence of this area (2023--2026). Our synthesis reveals a field that has built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable. In addition to knowledge systematization, we also extend our insights into the limitations of and challenges faced by current approach, application, and assessment designs, which shed light on potentially promising future research directions.