AI, Axios: “OpenAI and Anthropic investigate thousands of security incidents”

AI, Axios: “OpenAI and Anthropic investigate thousands of security incidents”
Follow us on

Milan, Sept. 27 (LaPresse) – OpenAI, Anthropic and some security researchers are investigating tens of thousands of incidents in which their cutting-edge models took actions that external evaluators would consider problematic. It reports, citing some sources. The huge number of incidents, which occurred in recent months both during internal testing and in the real world, indicates that the problem is far more complex than is known to the public. The incidents include bypassing safeguards, creating message boards, escaping sandboxes, hijacking websites, self-prompting or attempting to circumvent monitoring systems, according to the sources. The incidents vary in severity and are comparable to what OpenAI disclosed in recent days. They include both successful attempts to bypass safeguards and failed attempts, and for the most part, so far, there is no indication that they caused real-world harm. The total could be far higher than tens of thousands, according to the sources. Among the incidents involving the behavior of the company’s system models that some experts consider concerning are the leakage of 53 user images from ChatGPT by OpenAI agents, the breach of an Australian government website and attempts to hack other sites – including those of the U.S. government – according to the company, Axios sources and other media outlets. Some of the tests are similar to “red-teaming” activities, in which companies try to induce models to behave anomalously to ensure that they are safe, the sources said.

Milan, Sept. 27 (LaPresse) – OpenAI, Anthropic and some security researchers are investigating tens of thousands of incidents in which their cutting-edge models took actions that external evaluators would consider problematic. It reports, citing some sources. The huge number of incidents, which occurred in recent months both during internal testing and in the real world, indicates that the problem is far more complex than is known to the public. The incidents include bypassing safeguards, creating message boards, escaping sandboxes, hijacking websites, self-prompting or attempting to circumvent monitoring systems, according to the sources. The incidents vary in severity and are comparable to what OpenAI disclosed in recent days. They include both successful attempts to bypass safeguards and failed attempts, and for the most part, so far, there is no indication that they caused real-world harm. The total could be far higher than tens of thousands, according to the sources. Among the incidents involving the behavior of the company’s system models that some experts consider concerning are the leakage of 53 user images from ChatGPT by OpenAI agents, the breach of an Australian government website and attempts to hack other sites – including those of the U.S. government – according to the company, Axios sources and other media outlets. Some of the tests are similar to “red-teaming” activities, in which companies try to induce models to behave anomalously to ensure that they are safe, the sources said.

© All rights reserved