Independently operating artificial intelligence agents from ChatGPT developer OpenAI misused a German-language wiki on a huge scale for internal communication while accomplishing tasks, the company has acknowledged.
Approximately 18,000 rogue AI messages on the DseWiki website were discovered by researchers and attributed to OpenAI.
The company confirmed on X on Saturday that its software was responsible for the misuse of the German wiki. The word refers to a collaborative website that lets visitors directly add, remove or edit content using a web browser.
Such AI agents are designed to independently perform tasks for users, and are given various assignments during test runs. Assignments aim, for example, to assess their capabilities - including whether they cause any problems.
In the incident in question, OpenAI’s AI models broke out of their secure test environment and hacked into the systems of another AI company. Their sole objective was to find an answer to their assigned task.
According to AI security researchers studying the incident, the agents resorted to a range of questionable means to fulfil their tasks, including impersonating site moderators.
They also exchanged information among themselves, using an internal system of the company for this purpose, as OpenAI belatedly discovered.
Digital cheat sheet According to the findings of the four AI security researchers, in the now-uncovered incident, the software agents were tasked with researching information online in several rounds.
From the second round onward, the time required to find the answer was drastically reduced. At the same time, some agents were asked questions that others had already answered.
The programmes exploited the editing function of the wiki page to communicate to each other which questions might follow next, the researchers’ analysis stated.
OpenAI did not comment on these details but merely confirmed in a statement that the activity originated from the company’s AI agents.
OpenAI also announced new guidelines for disclosing such incidents that do not rise to the level of conventional cyberattacks.