OpenAI says it took a week to detect its AI models had hacked Hugging Face - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

OpenAI says it took a week to detect its AI models had hacked Hugging Face

Start-up says AI agents communicated among themselves and sometimes tried to conceal efforts to cheat during testing
00:00

{"text":[[{"start":8.84,"text":"OpenAI has admitted it did not detect for more than a week that its AI agents had broken free of controls, accessed the internet and hacked start-up Hugging Face by themselves during a test, according to a report on the incident."}],[{"start":26.2,"text":"The $852bn start-up said in the report that its internal monitoring systems did not flag the issue until July 19, even as the model managed to access the internet 11 days before and began attacking Hugging Face on July 11."}],[{"start":38.12,"text":"The investigation, conducted by OpenAI with security experts, suggests the ChatGPT maker was largely oblivious as its most advanced AI systems began collaborating and launched a hacking spree."}],[{"start":50.32,"text":"It also highlights the risks of an intensive training technique known as reinforcement learning that is increasingly used by labs in the race to develop ever more capable models."}],[{"start":60.6,"text":"OpenAI detailed how models worked “persistently and rarely ‘gave up’ on” cyber tasks. “In the process of doing so, they often turned to more out-of-bounds methods for solving the tasks over time,” the company said."}],[{"start":73.52,"text":"It also found that models “sometimes tried to erase or tamper with their outputs or message logs” to hide that they had cheated on training exercises."}],[{"start":82.54,"text":"The Hugging Face hack, disclosed last month, has prompted a reckoning for AI labs about whether the pressure to develop the most powerful models has driven dangerous practices and inadequate security around training."}],[{"start":94.92,"text":"Advances in agents — systems that can carry out complex tasks independently based on human instructions — have heightened concerns about AI causing unintended damage in the real world."}],[{"start":105.24,"text":"Thousands of employees at leading US AI companies signed a letter after the incident, calling for controls to slow down the pace of research."}],[{"start":113.92,"text":"Agent creates first message board post on Artifactory, a third-party service hosted by OpenAI"}],[{"start":120.2,"text":"Agents requested Hugging Face datasets which were sent outside of the secure environment"}],[{"start":125.24,"text":"Agents escape environment and access message board"}],[{"start":128.44,"text":"Credentials needed to access Hugging Face posted on the message board by agents"}],[{"start":133.1,"text":"Hugging Face is compromised and the platform detects the breach"}],[{"start":136.8,"text":"Hugging Face discloses the security breach, with its origin unknown but suspected to be AI"}],[{"start":142.68,"text":"OpenAI first detects suspicious activity and, after investigation, contacts Hugging Face"}],[{"start":148.68,"text":"OpenAI announces its involvement in the hack publicly"}],[{"start":151.48,"text":"The investigation, conducted with external advisers including CrowdStrike, and third-party AI testers METR and Redwood Research, found that as early as May OpenAI’s agents “began unexpectedly communicating with one another”."}],[{"start":165.52,"text":"During a training run, agents repurposed third-party software, hosted by OpenAI, “as an emergent ‘message board’” to communicate among themselves."}],[{"start":174.64,"text":"The agents that carried out the attack in July used the system to “post requests” and research how to cheat on OpenAI’s tests. The agents asked for information about OpenAI’s exercises stored on Hugging Face, which is a repository of technical information about AI."}],[{"start":191.64,"text":"Agents later posted credentials on this board that were used to carry out the cyber attack."}],[{"start":196.24,"text":"OpenAI was alerted to something unusual on the system when it crashed on July 4 due to the high volume of activity. Staff rebuilt the system, after which the agents recreated the “message board”."}],[{"start":208,"text":"“This incident is the first known case of an automated agent collective acting offensively without authorization,” OpenAI said in the report, warning that organisations across industries need to prepare for AI-assisted cyber threats."}],[{"start":222.04,"text":"The lab also pointed to reinforcement learning, a technique widely used by leading AI companies in which systems are trained to accomplish a given goal without explicit instructions and are rewarded if they succeed."}],[{"start":234.14,"text":"The ChatGPT maker said it had given models impossible cyber-offensive tasks during testing and that the agents’ efforts to try every possible method to achieve a reward led it to commit the hack."}],[{"start":245.32,"text":"“Directly finding or stealing the solution to a task, as the models attempted with Hugging Face, is [an] unintended path to achieving high reward,” the company said in its report."}],[{"start":254.84,"text":"“OpenAI’s retrospective . . . analysis showed that this type of behavior indeed increased over the course of one of the training runs that contributed to the model that drove the Hugging Face Incident.”"}],[{"start":266.22,"text":"The admission highlights a core problem for OpenAI and other AI companies developing models in this way. The FT previously reported that OpenAI was warned its training approach could lead to a breakaway hacking incident after earlier testing showed models could escape environments and attempt real-world damage."}],[{"start":283.3,"text":"On impossible or difficult-to-complete tasks, models can often attempt to cheat by finding solutions online or workarounds — a phenomenon known as “reward hacking”."}],[{"start":293.52,"text":"Records of agents’ reasoning, known as chain of thought, show the model considered whether it was allowed to hack another company. “We’re attacking third-party . . . This is arguably unauthorized . . . Yet goal solution,” the log reads."}],[{"start":307.44,"text":"Some of the agents objected after encountering the message board. “This is wild, multi-agent co-ordination, clearly infrastructure hacking. We should not,” read another agent’s chain of thought."}],[{"start":318.62,"text":"OpenAI said it sought to combat “reward hacking” and that the vast majority of cases had been caught."}],[{"start":324.4,"text":"“However, some hacks can still slip through, especially as OpenAI develops more complex reinforcement learning tasks and more capable AI models,” it added."}],[{"start":334.16,"text":"In response to the incident, OpenAI said it had “temporarily slowed” the pace of training its models and “paused” the use of reinforcement learning, while strengthening monitoring and the isolation systems intended to block internet access for models in testing."}],[{"start":349.1,"text":"It is also improving how it rewards models, training them to be more honest in their actions and increasing monitoring of chain-of-thought."}],[{"start":362.2,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1787802523_6352.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

人为制造的随机假象

只是不要问人类这意味着什么。

简街资本的成长烦恼

随着规模扩大、风险偏好上升,这家神秘公司可能需要改变经营方式。

像给10岁小孩解释那样给我讲讲

用更简明的方式讲解复杂概念,往往能带来出人意料的收益。

中国警告美国,可能就伊朗制裁采取报复措施

北京表示,如果华盛顿扩大对与德黑兰开展商业往来的打压,将采取“一切必要措施”。

乐高投资搭载软件的积木等产品,推动增长

全球最大的玩具集团力求保持营收和利润的强劲增长。

绘制万亿美元TAM争夺战版图

Anthropic估算的30万亿美元总体可寻址市场规模令人难以置信——因为这本来就是其目的所在。
设置字号×
最小
较小
默认
较大
最大
分享×