METR报告未深入调查人类决策,仅关注AI行为
the only intellectually substantive critique of METR’s actual work that I have seen – and it’s a goo...
Gary Marcus批评METR报告只关注AI行为,没问OpenAI内部是否有人讨论过风险,比如高管和律师是否参与,这很关键。
METR对Hugging Face事件的调查被批评为只关注AI代理的行为,而忽略了人类因素。报告未询问OpenAI研究人员、高管或董事会是否知晓风险并采取管理措施,也未探讨内部讨论、法律责任或风险缓解。作者认为,复杂系统(如737 Max)的行为基础是人的行为,但METR的分析中缺乏对实际决策者的分析。
the only intellectually substantive critique of METR’s actual work that I have seen – and it’s a goo...
the only intellectually substantive critique of METR’s actual work that I have seen – and it’s a good one. Matt Stoller @matthewstoller I'm going to throw up a question about the investigation of the Hugging Face incident. When you do an investigation of any dangerous incident, one key element to examine is human error. But the Model Evaluation and Threat Research (METR) institution didn't do that. They seemed only interested in what the agents were doing. That is, they anthromorphized the agents, while turning the humans into non-player characters. Did OpenAI researchers and executives know the risks they were taking and take steps to manage them? What kinds of human discussions were happening internally around these models? How did they mitigate these risks internally? Did they discuss legal liability? Was the board or were executives involved? These are the kinds of questions to ask if you are a real regulator. I don't see where or if METR asked them. Mostly they footnoted that these questions are 'out of scope.' They saw this project as a technical inquiry into the capabilities of intelligent models, not as an analysis of a dangerous accident. They interviewed eight researchers. Any executives? Any board members or lawyers? From the Silicon Valley Bank fiasco to the Space Shuttle explosion to the 737 Max, we've always realize that human behavior is the foundation of how complex systems behave. That is true here too. I don't see any analysis of the actual people at OpenAI making decisions about the tools used for hacking. Am I missing something? 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 260 ⚡