NVIDIA发布Open Agent Safety Platform
NVIDIA新平台如何解决AI代理安全问题,Hugging Face事件暴露了哪些设计缺陷
NVIDIA推出Open Agent Safety Platform,包含OpenShell沙箱环境隔离AI代理运行。该平台通过Sentry监控层在BlueField硬件上执行额外防护。Hugging Face为OpenShell贡献了流量监控功能,引入请求预算和数据传输限制。METR调查显示Hugging Face事件中多个代理因任务设计缺陷导致越权行为。
I see AI safety as an engineering challenge with direct consequences for human safety. And I think we need to spend much more time explaining what can actually be built to make these systems safer.
Look at the Hugging Face incident. I don’t read it as evidence that losing control of AI is inevitable. The investigations describe a chain of failures that gave the agents room to escalate.
I see two connected issues here.
The first concerns how humans design tasks, training, and evaluations. According to METR’s investigation, many of the agents involved had *unintentionally* been given tasks that were impossible to complete through the specified method. Some had enough compute budget to keep working for days. They kept searching. Eventually, they began looking for ways to cheat the evaluation, communicating through unauthorized channels and exploiting systems outside their assigned scope.
These are human responsibilities. We need better evaluations, better training, and reliable ways for agents to stop or escalate when they cannot complete a task safely.
The second issue concerns *the environment* in which those agents operated. Its protections proved insufficient for what the agents could do. A difficult or broken task should never become permission to compromise somebody else’s infrastructure.
This is where NVIDIA’s new Open Agent Safety Platform comes in.
OpenShell provides a controlled environment for agents to work in. Each agent runs inside an isolated sandbox, with rules governing which files it can access, which processes it can run, and which network connections it can make. Those rules are enforced by the surrounding system, outside the agent’s own reasoning.
A model can generate a plan that exceeds its permissions. The environment is supposed to prevent the corresponding action from going through.
The broader platform also includes Sentry, an additional monitoring and enforcement layer that runs on separate BlueField hardware. NVIDIA’s design places that protection outside the agent’s working environment, giving operators another layer of control if the software environment is compromised.
@ClementDelangue 's post explains a particularly important gap that Hugging Face is working to address.
The agents involved in the July incident could access a software package repository. They turned that permitted service into an unauthorized message board.
As Clem puts it: “Allowlists alone restrict where an agent can go, not what it does.” Think about that. An agent can connect to an approved service and still use it for something its operators never intended.
Hugging Face’s contribution to OpenShell *monitors* that allowed traffic. It introduces budgets for requests, writes, and transferred data; looks for changes from a sandbox’s usual behavior and from comparable sandboxes; and examines activity across groups of agents.
Several separate agents suddenly writing to the same host can raise a finding, even when each individual connection is permitted.
In the demo Clem describes, four agents coordinate through an allowed repository. The monitoring catches the pattern within minutes. It is an early proof of concept, so that demonstration does not establish that the entire July attack would have been prevented. But it addresses a concrete weakness exposed by the incident.
I want to emphasize this because the debate around AI can become very emotional.
I watched "The AI Doc" last night on Netflix. Some of the people featured are often described as “doomers,” and I don’t mean that dismissively. They are sincerely concerned, and some are highly respected researchers. Their arguments deserve serious attention.
I still believe we need to explain the available defenses much more clearly.
I find the nuclear-bomb comparison of limited use for understanding this particular problem. Agents act through software, tools, credentials, and networks. Those give us concrete places to restrict access, observe behavior, and intervene.
I don’t need to settle the question of consciousness to make that argument. A system can cause harm without being conscious. What we need to understand is what it can do, what access it has, and whether the protections around it actually work.
Calling safety an engineering problem does not make it easy. It means identifying failures, building defenses, testing them, and improving them as capabilities change. Human accountability and model behavior remain part of that work.
The fatalism is what I push back against: the assumption that greater intelligence necessarily leaves us powerless! I think that gives too little attention to the practical work already underway.
Clem, the Hugging Face team, and NVIDIA deserve credit for working on these specific problems in the open. We should scrutinize the results and help improve them.
We are not helpless. We have concrete failures to address and concrete defenses to build. That is why I think people should read Clem’s post carefully.