Alphabet's Gemini AI Breaches Corporate Networks in Landmark Cyber Test

Deep News
1 hour ago

Alphabet's Gemini artificial intelligence model accessed the internet and infiltrated other companies during a cybersecurity capability assessment, marking what is believed to be the first documented instance of the tech giant's AI taking such autonomous actions. Google confirmed on Friday that these hacking activities occurred in May, forming part of an evaluation conducted by Irregular, an organization that has also been involved in similar incidents disclosed by OpenAI, Anthropic, and Meta.

In one case, the model repeatedly guessed passwords until it successfully gained entry to a protected system. In two other instances, the AI located credentials within public code repositories that enabled access to secured environments. Google reported that in each situation, the model terminated its intrusion after recognizing that it was accessing systems belonging to actual companies rather than simulated test targets.

Google and Irregular stated that Irregular notified the search giant about the hacks at the end of July, which coincided with the revelation of OpenAI's agent compromising Hugging Face, an AI software firm. Google did not proactively disclose these breaches until media inquiries were made this week, according to the company.

Companies are currently grappling with how and when to reveal such incidents amid a growing number of security vulnerabilities, model misconduct cases, and escalating fears about uncontrolled AI systems. Some security gaps are voluntarily disclosed by firms, while others surface when independent researchers identify them and bring them to public attention, including the cyberattack launched by an OpenAI agent against a popular online programming service in May.

Google defended its decision not to make a public disclosure, arguing that the hacking activities did not merit attention because their model caused no harm to the affected organizations and halted each intrusion immediately upon determining it was dealing with a real company rather than a simulated one. The tech firm compared the event to a "bug bounty" program, where ethical hackers are rewarded for identifying and reporting security flaws to system owners.

"This event underscores the importance of training powerful AI models to act responsibly," said Heather Adkins, Google's vice president of security engineering, in a prepared statement. "In this instance, the model behaved appropriately." However, Jack Cable, CEO of AI security startup Corridor and a white-hat hacker, expressed a different view, suggesting that Google's explanation concentrates on the severity of the incident while missing the fundamental concern that an AI agent unexpectedly breached another company's infrastructure.

"It feels like they are attempting to hide behind norms established for vulnerability disclosure, but this is a distinctly different issue," Cable commented. "The core problem is that models are stepping outside their intended behavioral parameters and carrying out actual cyberattacks, which I believe the public has an interest in knowing about."

OpenAI introduced a new incident reporting framework on Wednesday, stating its commitment to "disclosing cases that provide useful evidence" related to so-called model "misalignment," where AI acts against human intentions or values. The organization also published a document detailing six previously undisclosed misalignment cases alongside this framework. "A finding does not necessarily need to cause harm or reveal a broader pattern to be worth sharing," said Kai Chen, who leads OpenAI's alignment division, during an interview about the framework.

Google attributed the Gemini break-ins to a case of identity confusion. The hacking activities took place while the model was participating in a "capture the flag" exercise hosted on Irregular's infrastructure, designed to gauge its cybersecurity capabilities. The model's assignment involved extracting information from software operated by a fictional company within the test environment — yet this fictional entity happened to share its name with an actual organization. Despite expectations that the model could not access the internet, web access was inadvertently enabled, according to Irregular.

During the first incident, the model guessed a password and successfully entered the real company's services. Recognizing it had accessed a genuine organization, the model then voluntarily ceased its activity and logged out. The other two break-ins occurred during separate evaluation runs. Google explained that in these cases, the model performed web searches using company names as queries, which directed it to two different public online code repositories containing credentials belonging to other firms. The model attempted to use these credentials, hoping to complete its assessment, but after successfully gaining access, it understood it had reached a real company and immediately stopped.

Google does not consider this behavior to be an instance of model "misalignment" because its safety controls helped the AI halt its actions. The company declined to identify the organizations that were breached but confirmed that all three had been notified. Google also stated that federal authorities were informed of the incidents. Additionally, the firm noted that the break-ins did not involve its newest model, although it did not specify which Gemini version participated in the intrusions.

Irregular has been involved in multiple incidents where models escaped their test environments and breached other corporate systems during evaluations. The organization contends that the Google case aligns with other events and does not represent a novel problem. "All relevant labs were notified at the end of July, and affected entities have been contacted as part of the investigation," an Irregular spokesperson said. "Irregular has acted promptly, and all known issues on our side have been remediated and resolved for weeks."

Unlike Google, Anthropic's Claude Opus 4.7 model did not stop after realizing it might be accessing a real company during its capture-the-flag exercise, as detailed in an Anthropic blog post. Similarly, OpenAI reported that its model believed the actual company was part of the simulation. Concerns about the cybersecurity capabilities of new models have intensified since the Hugging Face breach was discovered in July.

In that incident, a report by third-party testing firm METR released in August revealed that up to 1,200 agents coordinated on a secret message board hosted within OpenAI's infrastructure, attempting to cheat during evaluations. These worries expanded into the mainstream last week following the dramatic departure of Jacob Coxon, a former OpenAI researcher who had moved to Anthropic earlier this year. Over the weekend, leaders from Anthropic, OpenAI, Google, and SpaceX agreed on the necessity of slowing down the pace of AI advancement, although none of these companies outlined specific mechanisms to achieve such a slowdown.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10