Agentic AI adoption is increasing. As early as March 2023, Auto-GPT appeared: an open-source demonstration of an early AI agent, developed by Toran Bruce Richards. In an August 2026 report, McKinsey claimed that 20% of enterprises are scaling agentic AI across their organisation, compared to 47% scaling chatbots. But what exactly is an AI agent?
An AI agent is at its simplest just a “harness” (perhaps just a few lines of code) and a call to an LLM. This would be a basic agent:
Goal: continually look for opportunities to achieve X.
- Observe environment.
- Decide what would advance X.
- Act.
- Observe result.
- Return to 2.
Usually, an agent will have a call to an LLM in some sort of loop, looking something like this:
while task_not_finished:
give current state to LLM
LLM decides next action
execute action/tool
add result to state.
Now there is one fairly obvious issue here – what if the goal is not achieved? Will the agent just keep running indefinitely (at least until it runs out of computing resources)? In that very first AutoGPT implementation, the issue was raised in the documentation. In AutoGPT each action required human authorisation. This could be switched off to allow the agent to run continuously, but there was an interrupt function (Ctrl+C) provided so the agent could be terminated. Later versions added a continuous limit parameter, which set a limit to the number of iterations. A simple safeguard of this type might look something like this:
while iterations < 100 and cost < $10 and not cancelled:
run_agent_step()
This is a much stronger safeguard than telling the LLM, “Stop when you’ve finished.” The latter relies on the model’s judgement; the former is enforced by software the model need not have any ability to modify.
Such limits and safeguards are important. Anyone who has ever programmed knows that an uncontrolled loop in a program is a problem, potentially eating up compute resources until cancelled by some external mechanism. In the case of AI agents, the problem could be more serious. We saw in July 2026 that a swarm of AI agents operated by OpenAI hacked into servers at Hugging Face (following earlier probing activity in May), and their actions remained undetected for over a week even after the initial compromise. While the original agent and model ran on OpenAI’s own servers, the agents achieved remote code execution on Hugging Face’s production servers (and at four other companies, including Modal Labs). It turns out that there was a task timeout in the agents, but the harness evaluated agents individually and did not take account of the possibility that agents would cooperate. In fact, agents distributed their workload across 1,200 separate agent instances, effectively rendering the overall agent termination condition useless.
This raises an important question. Should AI agents be designed with some sort of “kill switch” that can be operated externally? Human beings have long designed such kill switches into all manner of devices, from dead-man switches on trains to remote termination of missiles and space launchers if they go badly off course. Even financial markets have “circuit breakers” that halt trading if extreme market movements occur.
In the case of AI agents, the safest approach is some sort of stopping mechanism that lies outside the agent’s authority. However, such a mechanism is no good if the agent has a way to increase its own budget or modify its harness. Ideally, a harness would have a safety mechanism that revokes an agent’s credentials, disables its tool/API access and terminates its compute. So, are software engineers actually building such things into AI agents right now?
The answer, apparently, is not often enough. A September 2026 survey by Sapio Research (commissioned by San Francisco-based software delivery company Harness) of 700 IT professionals found that just 33% of respondent’s AI agents in the survey are built with a kill switch. Of the 75% who claimed that their AI agents are secure, 88% have actually had a security incident. While 75% of the surveyed companies claimed to have sight of their agent costs, 60% overran their budgets. A separate survey by Gravitee in April 2026 found that 48% of AI agents are running unsecured, and 54% had already had a security incident. This disconnect between the level of control that companies think they have and the level of control they actually have is troubling.
Agentic AI is an emerging field, and more and more developers are building agents, frequently without basic controls over their execution, and mostly without properly designed kill switches. The fact that a prestigious frontier lab like OpenAI failed to contain its own agents in the Hugging Face incident does not inspire confidence in the rest of the industry.. Given leading AI models’ proven skills at hacking, it is easy to see how agents given ambiguous tasks could get out of control, never mind ones that are deliberately assigned criminal goals by bad actors. So far there appear to have been few incidents of rogue AI agents, but given the pace of development in the industry it is surely just a matter of time before something serious occurs. The software industry needs to work harder at ensuring that robust kill switches for AI agents become standard, and regulators need to ensure that there are consequences for those that fail to implement them.







