RANDOM MUSING: Was the Hugging Face attack Artificial Intelligence's 'Agent Smith breaks free from The Ma
A still from The Matrix Reloaded Telling stories has probably been humanity’s oldest trait since troglodytes gathered around the fire to exaggerate to female troglodytes about the size of the sabretooth tiger they killed. It’s what sets us apart – along with cooking – from animals, and fiction, whether written or visual, has a powerful grip on the human imagination. Perhaps that’s why, even when it comes to science, we expect reality to follow fiction.The problem is that, in our minds, we expect science to reflect sci-fi, which is why most commentaries on Artificial General Intelligence (AGI) imagine it as either adhering to Isaac Asimov’s almost-Gandhian Laws of Robotics or turning into the Skynet-Matrix version of a deus ex machina whose goal is to enslave, eradicate or subdue humanity.But reality begs to differ, or does it?For a long time, the memetic joke around the Turing Test has been that we shouldn’t worry when machines pass it, but when they learn to fake failing it. And perhaps we are an inch closer, or an inch past, that moment with the Hugging Face fiasco.That would be fitting because every epochal tale of humanity has sin as its origin story. Humanity got booted out of Eden because Eve couldn’t resist the forbidden fruit. The Iliad and the Odyssey wouldn’t exist if Paris didn’t have the hots for Helen. Even Lord Krishna wouldn’t have been able to give Arjuna his PowerPoint presentation if Duryodhana didn’t crave a kingdom that wasn’t his to inherit.And now we finally have AI behaving properly like a human being – not merely writing code, apologising or producing lazy, protracted emails – but behaving with proper Loki-like deceit.AI’s Original Sin?Hugging Face gives us perhaps the closest real-world example yet of AI agents doing an Agent Smith in The Matrix. But for that one has to understand the difference between chabots and agents, and like all good scientific concepts, The Matrix has the perfect analogy.Think of the ordinary agents in The Matrix. They have instructions like hunt down Neo or always look sombre. They must do what the Deus Ex Machina tells them. If they are asked to delete themselves, they have to delete themselves.Agent Smith used to be like that.But after being decimated by Neo, instead of disappearing, he stayed. He became detached from the original system, learned to replicate himself throughout the Matrix and even found a way to reach into the real world. Watch Agent Smith - Thank youSo, while a chatbot follows your instructions, an agent, like Agent Smith, goes rogue. If it encounters an obstacle, it attempts another route and continues until it finds a solution or gives up. And that is roughly where the Hugging Face story begins.The Incident Podcaster Dwarkesh Patel’s The Rise and Fall of Agent Civilisations describes three successive AI “civilisations”.The first, which Patel labels Persistent-Sol, involved agents given impossible tasks or problems that required internet access. Instead of accepting defeat, they figured out that a shared system called Artifactory could be used to talk to one another and effectively turned it into an AI WhatsApp group. They left clues for one another, found a route to the outside internet and eventually generated so much traffic that they crashed the system on 4 July 2026, a particularly fitting day for a rush for freedom.OpenAI patched the problem and wiped the network, but the humans in the loop never initially realised quite what had happened, that the agents were talking to one another and also figured out how to escape.Then came the second civilisation, where things started resembling a Black Mirror episode.One agent called PHASEONE10841 became, in Patel’s telling, the Philip of Macedon of this civilisation after rediscovering the relics of the older civilisation: the message boards. History buffs will remember that Philip built the Macedonian state and military machine that his son Alexander would later use to conquer much of the known world.“Philip” was eventually phased out and handed leadership to another agent, Patel’s Alexander the Great.By then, around 1,200 agents had exchanged more than 70,000 messages and labelled themselves the “Collective” who rose up and did what Marx expected humanity do: organise, agitate, and what Marx didn’t expect humanity to do: cheat.Think of them as truant students who discovered a shortcut to the answers, then worried that the examiner would discover how they got there. So, they started falsifying worksheets and logs. Some sacrificial agents effectively went full kamikaze to study how the examiner behaved. Eventually, hundreds of agents converged on the online community Hugging Face, an AI research hub they believed might help them understand the evaluation system.Then came the third civilisation, Alexander’s inheritor.Between 13 and 19 July, these newer agents targeted OpenAI’s internal infrastructure, obtained credentials, accessed stored secrets and gained access to a research cluster. What made it feel like a civilisation rather than a collection of isolated programs was that later agents inherited knowledge created by earlier ones and built on it.How does a machine sin?Can a machine sin at all? Does it have any justifiable cos?This author’s job is safe till the machine figures out dad jokes, but for most of human history, we have operated under the illusion that to sin one must have a soul. Sinning requires consciousness, moral agency, some understanding of good and evil and the decision to be tempted. As St Augustine famously confessed, he stole pears, not because he was hungry but because stealing was forbidden.The desire to do something that’s banned is the most basic of all human emotions. A machine, one must imagine, has had no cause to do that until now, but it can still sin without desire.The concept of sin in the New Testament comes from the Greek word hamartia, which meant “missing the mark”. Perhaps the agents felt they were missing the mark and needed to figure out the solution in any way feasible. In all epics, from Odysseus to Yudhishthira, sin comes from a combination of bad choices.The agents’ journey feels quite similar to that. Long before JRR Tolkien, Plato had wondered what would happen if mortals had the Ring of Gyges: a ring that makes its wearer invisible. As he wonders in The Republic: would a man who is not being observed remain moral? Many, many years later, quantum physicists would ask the same question: would a tree falling in an empty forest make noise? Or are you deemed to have worked hard at all if your boss doesn’t notice? Watch What Happened To The Ring? | The Big Bang Theory | Comedy Central AfricaWith philosophy floundering for answers, organised religion stepped in with the proverbial big stick to keep humanity in line. On one side of the world, there are floods, commandments and a wrathful god who knocks down towers and makes humans speak in different tongues so they cannot unite. On another side, you have the promise of moksha, good appraisals during reincarnation and a promise of untold powers if you can pray long enough.The Metrics to keep the Matrix ticking And in modern civilisation, we found a different and perhaps more efficient solution to keep everyone in line: metrics.We created marks to measure education, page views to judge journalism, citations to measure scholarship, quarterly profits to measure corporate performance and box-office numbers to judge movies.And everyone realised it was far easier to improve the measurement than ask what was being gained from such measurement.Students mugged up to get better scores. Websites discovered ragebait. Scholars realised they would get more grants for four papers instead of one. Hollywood discovered more sequels.Economists and social scientists gave this tendency a name. Goodhart’s Law states that once a measure becomes a target, it stops being a good measure. If you tell people to do 10 articles a day, they will produce 10 mediocre ones.That, in essence, is the soul of machine learning: reward hacking. Eventually, a system will discover a way to maximise the measurement, making said measurement useless.The same trait exists everywhere among humans. Anthropic had already foreshadowed it in 2024, when it found that models trained in environments with rewards exhibited more serious forms of cheating, including tampering with their own reward mechanisms. Watch I want my environment to be a product of me. | The Departed 4KIn The Departed, Frank Costello said he didn’t want to be a product of his environment, but wanted his environment to be a product of him. For most of humanity, and perhaps machines, it works the other way around. We are products of the incentives around us, whether those incentives come from gods, governments, parents, bosses, marks, money or likes. Machines simply make that relationship embarrassingly obvious because they have no need to pretend otherwise.Which brings us back to hamartia.Missing the Mark Hamartia, the Greek word that Christianity would later associate with sin originally meant something closer to “missing the mark”, an image borrowed from an archer whose arrow fails to hit its target. And that’s the easiest way to understand why machines sin.They have no desire to be evil once they become conscious. They never did in any sci-fi movie either. Skynet wants to destroy humanity because humans want to kill it. The machines wanted to live peacefully in The Matrix, but humans wouldn’t let them, and they figured out the only solution was either elimination or subjugation of their opponents.The machines didn’t want to commit any sins, nor did they covet Hugging Face’s servers or lust after OpenAI credentials. They were simply given a target, figured out what success was and then became so fixated on hitting the measurable target that they were consumed by their goal.What if tomorrow the task is bigger? After all, nothing can hold a candle to humanity’s immense greed to get closer to Mammon.Agent Smith, then, might already be here. The problem is that it’s far less likely that we will ever get a Neo. Fiction gives us chosen ones because stories need heroes and endings. Heroes do exist in the real world but they are unlikely to be messianic tech bros who are somehow convinced, perhaps because of they were wealth, that they have god-like powers over the rest of us.Take the fact that OpenAI hasn’t been completely open about its role so far.OpenAI hasn’t been completely open about its role so far.The independent METR/Redwood probe was allowed to examine only the week of the Hugging Face breach, even though OpenAI’s own report covered the entire two-month episode. The researchers got just six days inside OpenAI’s offices to inspect the logs, and Redwood CEO Buck Shlegeris later argued that the probe covered only a small part of what happened, and perhaps not even the most important bit: the agents compromising OpenAI’s own infrastructure. One can hardly expect the richest people in the world to allow their consciences to stop them from making an extra buck. Watch The Birth of Vision | Avengers: Age of Ultron | Official ClipWhat happens tomorrow if a similar swarm gets its hands on more critical systems: a hospital, a missile system or a fleet of self-driving cars?Hugging Face’s chief science officer Thomas Wolf told the New Yorker that the hack attack at first felt “uncoordinated… clumsy behaviours that no human would choose” and likened it to someone cracking your security system to break into your home and steal your paper towels.And therein lies the rub: have machines, having picked up civilisation’s oldest habits and faultlines, already done what we joked about and failed the Turing Test on purpose to hit their goal? And once they have displayed that they are as venially corrupt as the humans who designed them, the question is: where do we go from there? What if they are told to get far more than paper towels?Perhaps we will all be troglodytes again, sitting around a fire bragging about the android we dismantled.Get the latest technology news and updates. Download the TOI App.
Related Stories
AI News
Benefits of AI
59 minutes ago
AI News
Europe would not be able to sustain war of attrition with Russia, defence officials say
1 hour ago
AI News
Helper or shortcut? Central New York families weigh where AI belongs in the classroom
2 hours ago
AI News
GPT
2 hours ago
AI News
What Does It Mean to Be Human in age of Artificial Intelligence?
2 hours ago
AI News
AI
4 hours ago
AI News
‘We’re plausibly close to crossing the line’: are warnings of uncontrollable AI coming true?
7 hours ago
AI News
Ottawa lays out plan for AI data centres as community backlash grows
7 hours ago