It also exposed a second, less obvious challenge.
During the investigation, Hugging Face reportedly found that some hosted frontier models struggled to assist with forensic analysis because their safety guardrails restricted their ability to analyse real attack commands, exploit payloads and other security artefacts. In effect, the models could not consistently distinguish between authorised defensive investigation and offensive cyber activity.
To support aspects of the response, Hugging Face instead turned to GLM-5.2, an open-weight model developed by Z.ai and operated locally within its own environment.
This is the AI resilience conundrum. As models become more capable, frontier firms face an increasingly difficult challenge of preventing systems from pursuing unintended pathways to achieve their objectives. For enterprises, the controls designed to make advanced AI safer may also constrain legitimate defensive action.
The safeguards designed to make advanced AI systems safer are essential. Yet organisations must also consider what happens when those same safeguards limit their ability to investigate, contain or respond to a genuine incident. As AI becomes increasingly embedded within critical business operations, resilience will be defined by an organisation’s capacity to adapt and respond when controls, models or platforms are tested. The challenge is no longer simply keeping AI contained. It is ensuring that when containment fails, organisations still retain control.
So, when AI behaves beyond its intended boundaries, how does an organisation respond if the controls designed to improve safety also constrain legitimate defensive capability?
Why responsible AI must extend beyond the model
The answer lies in broadening how organisations think about responsible AI5. Guardrails remain essential. Their purpose is to reduce the risk that advanced capabilities are used to cause harm. The challenge is that specialised defensive activity can resemble offensive activity when assessed only through prompts, commands or technical artefacts. Context matters: who is asking, what authority they hold, which environment is being investigated and what oversight surrounds the activity.
As AI becomes more agentic, responsible AI must extend beyond the model itself. An AI system includes the tools it can invoke, the permissions it holds, the data it can access, the infrastructure on which it operates and the people accountable for its decisions.
Organisations need to combine pre-deployment testing with runtime controls. Models should operate with least-privilege access, defined network boundaries, observable behaviour and intervention points that allow authorised personnel to interrupt or isolate activity when necessary. Human oversight must be technically enforceable and operationally understood, not simply documented in policy.
The objective is not to weaken guardrails. Neither is it to assume that open-weight or locally operated models are inherently safer. Every model and deployment approach carries its own security, governance, performance and resilience considerations. The goal is to establish layered, context-aware controls that support legitimate use while preserving safety. Done well, this does not slow AI adoption. It creates the confidence required to deploy increasingly capable systems in increasingly consequential settings.
Four questions for boards and governments
If AI resilience is ultimately about maintaining control when systems, safeguards or platforms are tested, four questions deserve particular attention. We call this the AI resilience lens, a framework built around four interconnected dimensions: Control, Optionality, Continuity and Sovereignty. Together, they provide a practical way for boards and governments to assess whether they can retain oversight, resilience and operational control as AI becomes embedded in critical activities.
Control
Can AI remain within its intended operating boundaries? When behaviour deviates, can the organisation detect it quickly, understand what is happening and intervene before disruption spreads? Resilience depends on preserving human decision-making authority, maintaining visibility of AI-driven activity and ensuring systems can be paused, isolated or recovered when necessary.
Optionality
How dependent is the organisation on a single model, provider or ecosystem? AI strategies should consider a mix of hosted, open-weight and locally operated models, with trusted alternatives assessed against consistent standards for capability, security, responsible AI and resilience. The objective is to create meaningful choice while avoiding unnecessary complexity or new concentrations of risk.
Continuity
Can critical services continue if a preferred model becomes unavailable, behaves unpredictably or is unable to provide legitimate assistance? Leaders should identify which AI-enabled capabilities are essential, determine acceptable levels of service degradation and regularly test how operations would continue under different failure scenarios.
Sovereignty
As discussions around sovereign AI gather pace, the focus must extend beyond where a model was developed or where data is stored. The more important question is who retains operational control during a crisis and whether critical capabilities can continue if access to a platform, provider or broader ecosystem is disrupted. Sovereignty is ultimately a question of control across data, infrastructure, models, operations and decision-making. Local data hosting alone does not eliminate dependence on external model policies, software updates, connectivity or specialist expertise. Resilience may depend just as much on access to trusted alternatives when primary capabilities are constrained.
From model assurance to ecosystem resilience
The AI resilience lens highlights a broader shift in how organisations must think about AI risk. These are no longer purely technical questions. They go to operational resilience, strategic autonomy and national competitiveness. As AI becomes embedded across software development, security operations, customer service, infrastructure management and decision-making, failure may no longer remain confined to a single model or technology stack.
It can propagate through platforms, data pipelines, suppliers and operational processes. This builds on the lesson of Project Glasswing: risk has not fundamentally changed, but the speed at which exposure becomes impact has accelerated. Just as COVID exposed hidden dependencies across physical supply chains, AI is beginning to expose hidden dependencies across digital ones, including models, platforms, infrastructure, data, skills and third parties.