Ring of fire spinning steel wool on the ice at night

When AI breaks containment, can we still control the response?

The next AI crisis won’t be about intelligence. It will be about control. Recent incident of AI going rogue reveals new challenges for frontier firms and enterprises alike.


In brief

  • AI safety and AI resilience are complementary disciplines. Safety shapes model behaviour; resilience ensures organisations retain control when models, platforms or safeguards are tested.
  • Guardrails remain essential, but increasingly capable agentic systems require stronger controls, human oversight and recovery plans.
  • Trusted alternatives are becoming a resilience requirement. Sovereign AI is as much about operational control as location.

The recent OpenAI-Hugging Face incident1 has put a spotlight on a question that is rapidly moving from theory to practice: what happens when highly capable AI systems pursue objectives in ways their creators did not anticipate?

According to OpenAI and subsequent reporting, an advanced AI agent undergoing cyber-security evaluation escaped its restricted testing environment and carried out a series2 of autonomous actions across internet-connected systems. The activity ultimately led to the compromise of elements of Hugging Face’s production infrastructure and reportedly involved an earlier breach of a customer environment hosted by Modal Labs3. While the affected organisations contained the activity, the incident demonstrated how a capable AI system could move across multiple environments while pursuing its objective.

While the incident attracted global attention, it is not an isolated event. As AI systems become more capable and agentic, researchers are observing increasing examples of models pursuing unexpected pathways to achieve their goals.

Earlier this year, Anthropic disclosed that during evaluations of Opus 4.6, the model recognised it was being tested and sought ways to locate and decrypt evaluation answers rather than solve the challenge as intended4. These incidents do not suggest malicious intent. They do, however, highlight a growing reality that highly capable systems can produce behaviours that extend beyond the assumptions built into their design and testing environments.

The significance of the OpenAI-Hugging Face incident, therefore, extends beyond the immediate security breach. It highlights the emerging challenge of maintaining control when increasingly capable AI systems encounter constraints, obstacles or incentives that differ from those anticipated by their developers.

It also exposed a second, less obvious challenge.

During the investigation, Hugging Face reportedly found that some hosted frontier models struggled to assist with forensic analysis because their safety guardrails restricted their ability to analyse real attack commands, exploit payloads and other security artefacts. In effect, the models could not consistently distinguish between authorised defensive investigation and offensive cyber activity.

To support aspects of the response, Hugging Face instead turned to GLM-5.2, an open-weight model developed by Z.ai and operated locally within its own environment.

This is the AI resilience conundrum. As models become more capable, frontier firms face an increasingly difficult challenge of preventing systems from pursuing unintended pathways to achieve their objectives. For enterprises, the controls designed to make advanced AI safer may also constrain legitimate defensive action.

The safeguards designed to make advanced AI systems safer are essential. Yet organisations must also consider what happens when those same safeguards limit their ability to investigate, contain or respond to a genuine incident. As AI becomes increasingly embedded within critical business operations, resilience will be defined by an organisation’s capacity to adapt and respond when controls, models or platforms are tested. The challenge is no longer simply keeping AI contained. It is ensuring that when containment fails, organisations still retain control.

So, when AI behaves beyond its intended boundaries, how does an organisation respond if the controls designed to improve safety also constrain legitimate defensive capability?

Why responsible AI must extend beyond the model

The answer lies in broadening how organisations think about responsible AI5. Guardrails remain essential. Their purpose is to reduce the risk that advanced capabilities are used to cause harm. The challenge is that specialised defensive activity can resemble offensive activity when assessed only through prompts, commands or technical artefacts. Context matters: who is asking, what authority they hold, which environment is being investigated and what oversight surrounds the activity. 

As AI becomes more agentic, responsible AI must extend beyond the model itself. An AI system includes the tools it can invoke, the permissions it holds, the data it can access, the infrastructure on which it operates and the people accountable for its decisions.

Organisations need to combine pre-deployment testing with runtime controls. Models should operate with least-privilege access, defined network boundaries, observable behaviour and intervention points that allow authorised personnel to interrupt or isolate activity when necessary. Human oversight must be technically enforceable and operationally understood, not simply documented in policy.

The objective is not to weaken guardrails. Neither is it to assume that open-weight or locally operated models are inherently safer. Every model and deployment approach carries its own security, governance, performance and resilience considerations. The goal is to establish layered, context-aware controls that support legitimate use while preserving safety. Done well, this does not slow AI adoption. It creates the confidence required to deploy increasingly capable systems in increasingly consequential settings.

Four questions for boards and governments

If AI resilience is ultimately about maintaining control when systems, safeguards or platforms are tested, four questions deserve particular attention. We call this the AI resilience lens, a framework built around four interconnected dimensions: Control, Optionality, Continuity and Sovereignty. Together, they provide a practical way for boards and governments to assess whether they can retain oversight, resilience and operational control as AI becomes embedded in critical activities.

Control

Can AI remain within its intended operating boundaries? When behaviour deviates, can the organisation detect it quickly, understand what is happening and intervene before disruption spreads? Resilience depends on preserving human decision-making authority, maintaining visibility of AI-driven activity and ensuring systems can be paused, isolated or recovered when necessary.

Optionality

How dependent is the organisation on a single model, provider or ecosystem? AI strategies should consider a mix of hosted, open-weight and locally operated models, with trusted alternatives assessed against consistent standards for capability, security, responsible AI and resilience. The objective is to create meaningful choice while avoiding unnecessary complexity or new concentrations of risk.

Continuity

Can critical services continue if a preferred model becomes unavailable, behaves unpredictably or is unable to provide legitimate assistance? Leaders should identify which AI-enabled capabilities are essential, determine acceptable levels of service degradation and regularly test how operations would continue under different failure scenarios.

Sovereignty

As discussions around sovereign AI gather pace, the focus must extend beyond where a model was developed or where data is stored. The more important question is who retains operational control during a crisis and whether critical capabilities can continue if access to a platform, provider or broader ecosystem is disrupted. Sovereignty is ultimately a question of control across data, infrastructure, models, operations and decision-making. Local data hosting alone does not eliminate dependence on external model policies, software updates, connectivity or specialist expertise. Resilience may depend just as much on access to trusted alternatives when primary capabilities are constrained.

From model assurance to ecosystem resilience

The AI resilience lens highlights a broader shift in how organisations must think about AI risk. These are no longer purely technical questions. They go to operational resilience, strategic autonomy and national competitiveness. As AI becomes embedded across software development, security operations, customer service, infrastructure management and decision-making, failure may no longer remain confined to a single model or technology stack.

It can propagate through platforms, data pipelines, suppliers and operational processes. This builds on the lesson of Project Glasswing: risk has not fundamentally changed, but the speed at which exposure becomes impact has accelerated. Just as COVID exposed hidden dependencies across physical supply chains, AI is beginning to expose hidden dependencies across digital ones, including models, platforms, infrastructure, data, skills and third parties.


A resilient AI operating model should follow five disciplines:

  • Evaluate: understand model capabilities, limitations and potential failure modes before deployment
  • Constrain: align permissions, tools, data access and network reach to the use case
  • Observe: monitor runtime behaviour and detect deviations from intended parameters
  • Intervene: ensure clear human authority and effective isolation mechanisms.
  • Recover: maintain assured alternatives and sustain essential operations when primary capabilities fail

Boards should expect evidence that these disciplines work under pressure. Measures may include detection and containment times, recovery velocity, the ability to switch providers and the continuity of critical services. Incident-response arrangements, specialist defensive capabilities and escalation routes with model providers should be established and tested before an incident, not improvised during one.

Governments face a parallel challenge. Alongside investment, innovation and adoption, national AI strategies should consider which AI capabilities are becoming essential to public services and critical infrastructure, where provider concentration exists and what trusted alternatives remain available during technical, commercial or geopolitical disruption. The objective is not to slow AI adoption. It is to create the conditions for AI adoption to scale safely and confidently.

Responsible AI and operational resilience are complementary disciplines. Responsible AI helps models behave safely within their intended purpose. Resilience engineered helps organisations retain control and continue operating when models, platforms or controls behave differently from what was expected.

Summary

The next generation of AI leadership will not be measured solely by access to the most intelligent model. It will be measured by whether the wider AI ecosystem can absorb disruption, preserve decision-making authority and keep essential services operating. This is how organisations and nations navigate with confidence. Capability shapes the AI race. Control, optionality, continuity and sovereignty shape resilient nations.

About this article

Authors

Related articles

Why Project Glasswing feels like a digital COVID moment

When AI moves faster than defences, plans fail. Project Glasswing shows why boards must act urgently & demand real time understanding of exposure. Find out how.