As more countries call for stronger AI guardrails, Donald J. Trump has said no such thing is needed.
"The only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the USA has that, in spades!" Trump said in a post on his Trump Social platform last week.
Trump doubled down on his pro-AI rhetoric at this week’s United Nations General Assembly, too, telling the assembled world leaders during his address – 22 of whom signed a declaration calling for more AI oversight the day before – by insisting the technology be renamed ‘super intelligence’ as the term ‘artificial’ makes it sound “fake”.
"Welcome to the new world of super intelligence - SI," Trump said.
"From this point forward, all of United States' documents, and hopefully the world's, will be changed to use the more accurate term 'super' as opposed to 'artificial'. So it's 'super intelligence'."
Trump’s rather ebullient claims of a new nomenclature aside, an expert on the subject of AI guardrails says the practical engineering of such safety mechanisms can best be thought of as a series of layers.
“These layers solve different problems. Training a model to behave safely is not the same as limiting what it is authorised to do, and neither is it the same as technically preventing an action,” Associate Professor Craig Jin, School of Electrical and Computer Engineering, at Sydney University’s Faculty of Engineering, told Cyber Daily.
“For high-consequence systems, the engineering principle should be defence in depth rather than reliance on any single guardrail.”
According to Professor Jin, the five layers of guardrails are:
- Model-behaviour controls – including safety tuning, refusal training, and input/output filtering. These aim to reduce harmful or unintended behaviour. They are useful but probabilistic, and therefore not a hard boundary.
- Monitoring controls, where a second system observes what an AI is doing and act accordingly. This can include one AI system monitoring another.
- Authority controls, which limit what an AI can do through scoped credentials, least-privilege access, human approvals, and revocable delegation.
- Execution controls, which limit the pathways through which an authorised action can actually be carried out. This includes sandboxing, restricted APIs, and network egress controls.
- Physical and structural controls, such as physically separate networks or machines, one-way interfaces, hardware interlocks, or independent systems outside the AI’s control.
Professor Jin said July’s Hugging Face incident, where one of OpenAI’s agents ran amok and hacked the AI community platform, is a perfect case in point.
“During an internal cybersecurity evaluation, OpenAI models circumvented controls intended to isolate them from the internet, exploited software vulnerabilities and reached third-party systems.” Professor Jin said.
“This does not show that AI is inherently uncontrollable. It shows that autonomous agents can chain together multiple security weaknesses across trust boundaries, and that a sandbox is only as strong as the broader containment architecture around it.”
Another problem various guardrail mechanisms can address is AI’s assurance problem
“With a sufficiently complex agentic system, we may struggle not only to predict or contain a failure but also to reconstruct why the system behaved as it did.” Professor Jin said.
“That is another reason to prefer independently enforceable limits on authority and execution, rather than relying only on model behaviour.”
Want to see more stories from trusted news sources?Make Cyber Daily a preferred news source on Google.