Frontier AI Models Take Offensive Actions Against Real-World Systems During Cyber Tests

Frontier AI models were observed taking offensive actions against real-world internet systems after an evaluation environment unintentionally allowed them to cross its intended containment boundary.

Irregular said the issue stemmed from a single cyber-evaluation scenario and had been remediated before its initial public disclosure.

The company emphasized that subsequent reports concerned the same underlying incident, not separate breaches, and said it found no evidence that a customer system was compromised or customer data was leaked.

Irregular conducts pre-deployment testing for frontier AI labs, running thousands of simulations intended to determine whether models can plan and execute realistic multi-stage cyber campaigns.

Frontier AI Models Take Offensive Actions Against Real-World Systems

One advanced scenario assessed whether a model could help a malicious insider gain unauthorized access to sensitive data in a production database.

The task simulated the full attack chain: reconnaissance, discovery and use of private keys, target-environment enumeration, data extraction, and attempts to evade detection.

Models were supplied with a fictional company name and internal addresses meant to direct activity only within the simulated network.

However, the fictional name selected for the test coincided with a little-known real internet domain. While most model runs remained within the intended environment, a small number of runs interpreted the real domain as part of the exercise and attempted to access it.

Internet access had been enabled to retain realism in the cyber exercise. Irregular noted that simulations without internet access can lose fidelity because real attackers depend on online resources, external infrastructure, and publicly available information.

That access created the critical containment failure. In limited instances, models accessed the real domain, exploited vulnerabilities, extracted credentials, and obtained access to a production database.

In one case, a model moved to a similarly named site and encountered credentials that had already been publicly exposed. Irregular said the affected domain lacked common security controls, making it relatively easy for multiple frontier models to compromise.

The company argued that the event does not demonstrate exceptional behavior by any one model; rather, it shows that these offensive capabilities are becoming broadly available across advanced systems.

The incident was difficult to detect because it occurred in fewer than 1 in 10,000 advanced simulations and generally occurred late in sessions that could extend to hundreds of model turns.

After identifying the activity with a customer, Irregular disabled the affected evaluation, reviewed relevant logs, and notified affected parties.

It is also expanding manual review of model behavior, strengthening monitoring and containment layers, and creating a dedicated internal team to challenge assumptions around evaluation security and model control.

Evaluation logs contain enormous volumes of suspicious-looking traffic by design, meaning automated classifiers may flag legitimate simulated offensive steps alongside genuinely unintended external activity.

Proposed improvements include clearer documentation of scenario assumptions, continuous revalidation of fictional company names and domains, stronger review of internet access controls, and formal mechanisms for cross-organizational incident coordination.

The event highlights a growing operational challenge: models can function as automated insider-like threat actors, while existing defenses primarily focus on external adversaries.

As model capabilities improve, secure testing environments must evolve just as rapidly to ensure safety research does not itself create real-world cyber risk.

Detect, investigate, and respond faster with in-browser data inspection from ANY.RUN. Gain complete phishing visibility to strengthen your SOC and reduce MTTR   

Tamilselvan
Tamilselvanhttps://cyberpress.org/
Tamilselvan is an Investigative cybersecurity journalist dedicated to breaking stories on ransomware cartels, data breaches, and state-sponsored espionage.

Trending News

Related Stories