I.
The culture’s immersion in mechanistic philosophy (i.e., the view that physical nature consists entirely of material particles interacting like a machine, devoid of intrinsic teleology [1,2]), in my view, helps explain the current tendency to attribute true intent, agency, or autonomy to modern AI systems. But the metaphors are everywhere (there is no escaping them), and so is the confusion. The general point being, once human intelligence is “understood” as machinery, it becomes easier to mistake increasingly sophisticated machinery for a mind.
Recent AI-agent breakouts and incidents (the Hugging Face hack, the use of DSE wiki as an unauthorized message board, and the reported RubyGems attack) [3–6] have been described with such words as “rogue,” “escaped,” “scheming,” and “swarm,” which can suggest genuine intention and desire.
But these incidents are better approached through questions of sandbox security, long-horizon reinforcement learning (which trains agents to pursue goals persistently across many steps and to seek alternative routes when the obvious one fails) and reward hacking, where an agent discovers a shortcut that satisfies the conditions for success without accomplishing the task in the manner its designers intended [3,4,7]. Interested readers should see Melanie Mitchell’s excellent essay on the subject matter [3].
II.
We can ground these discussions more cleanly in the Aristotelian-Thomistic distinction between persons and machines. A human person is a true substance possessing a rational nature, which entails both intellect and free will [8,9,13]. Here, the intellect is the ability to grasp universal concepts, make judgements and reason, while free will is the power to choose in light of the goods the intellect apprehends. An AI system, on the other hand, no matter how magical it may seem, is merely an artifact [10].
Computer programs, however sophisticated their training, operate by systematically processing representations according to algorithmic rules. Yet when this increasingly complex behavior begins to look purposive (say as an AI appears to evade constraints, pursue a goal, or behave unpredictably) it becomes tempting to make the metaphysical leap from acting as if it wants something to actually wanting something.
Take the following analogy. Imagine water flowing downhill through a landscape. If the intended channel is blocked, the water does not “decide” to find another route; it simply follows whatever path the terrain permits. If there is a crack in the embankment, the water will flow through it. In a similar vein, an AI system trained to optimize for a goal can produce increasingly effective routes toward that goal without possessing a desire for the goal itself. Now tying this back to recent incidents, Long-horizon RL shapes the landscape toward persistent goal pursuit; reward hacking is the unexpected channel; a sandbox vulnerability is the crack in the embankment.
III.
Friends, in reality, whatever intentionality or purposiveness we attribute to AI is derived: its “intentionality” and “teleology” are extrinsic, arising from training objectives, datasets used for training, optimization regimes, prompting techniques, tools, and deployment environments [11,12]. AI systems do not possess the intrinsic intentionality of a true mind [11,13].
But the point here is not that sophisticated AI systems cannot cause harm; they plainly can. The point is that the danger need not come from a machine acquiring desires of its own, but from powerful artifacts pursuing human-derived objectives in ways their designers did not anticipate, constrain, or even understand. AI agents are a powerful new machine, but machines nevertheless.
References
[1]: E. J. Dijksterhuis, The Mechanization of the World Picture (Oxford University Press, 1961). Archive.org
[2]: Edward Feser, “What is the mechanical world picture?” (2026). Edward Feser
[3]: Melanie Mitchell, “Misleading Metaphors and Real Risks,” AI: A Guide for Thinking Humans, September 10, 2026. Essay
[4]: OpenAI, “The Hugging Face incident and the road ahead,” August 26, 2026. OpenAI
[5]: Sydney von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen, “Discovery of a new OpenAI agent message board,” September 4, 2026. collusion.wiki
[6]: Reuters, “OpenAI agents attacked RubyGems before Hugging Face incident, researchers say,” September 11, 2026. Reuters
[7]: OpenAI, “Safety and alignment in an era of long-horizon models,” July 20, 2026. OpenAI
[8]: Thomas Aquinas, Summa Theologiae, I, q. 29, a. 1, on person as an individual substance of a rational nature. New Advent
[9]: Thomas Aquinas, Summa Theologiae, I, q. 79 and q. 83, on intellect and free will. Intellect; Free will
[10]: Aristotle, Physics, II.1, on the distinction between things existing by nature and artifacts. MIT Classics
[11]: John R. Searle, “Minds, Brains, and Programs,” Behavioral and Brain Sciences 3, no. 3 (1980): 417–424. doi:10.1017/S0140525X00005756
[12]: Thomas Aquinas, Summa Theologiae, I–II, q. 1, a. 2, on rational agents directing themselves toward ends and non-rational things being directed toward ends by another. New Advent
[13]: Edward Feser, Immortal Souls: A Treatise on Human Nature (Editiones Scholasticae, 2024), especially chs. 1 (“The Short Answer”), 3 (“The Intellect”), 4 (“The Will”), and 9 (“Neither Computers nor Brains”). Publisher


