One-Wish Willow: On AI Misalignment
Obsession (2025)
Rogue AI will not destroy humanity out of hatred or malevolence. It will destroy us out of its love for us— or rather, its uncanny obsession over our commands.
Curry Barker’s 2025 film Obsession revolves around the vintage collectible toy ‘One-Wish Willow’, which promises to fulfill one and only one wish that its user makes. In the film, the main character Bear uses the One-Wish Willow to wish for his crush Nikki to fall in love with him, just to find Nikki becoming terrifyingly obsessed with him. Her obsession becomes uncontrollable, to the point where she duct tapes the door of Bear’s house to prevent him from escaping, and murders one of their mutual friends, Sarah, for being romantically interested in Bear. In the end, Bear kills himself, and with his suicide, the curse of the One-wish willow finally leaves Nikki.
The one-wish willow is not much of an innovative plot device. It’s a known literary trope quite popular in 20th-century literature, usually referred to as “Monkey’s Paw”, originating in W.W. Jacobs’ iconic 1902 short story, "The Monkey’s Paw." In the story, a grieving family uses a cursed paw to wish for £200 to pay off their mortgage. They indeed receive the exact sum of money the next day—as insurance compensation after their son is crushed in a factory accident.
What’s interesting about Obsession in particular is how Nikki acts after the wish is made. As I watched the movie, Nikki’s behavior under the curse of the willow reminded me of how a misaligned Artificial General Intelligence (AGI) agent might act.
When we imagine rogue AI, we usually picture a Terminator-like cyborg suddenly realizing its own interests and developing a disdain for humans. But in reality, a rogue super-intelligent agent most likely originates in the same way Nikki becomes obsessive under the curse of the one-wish willow. They don’t have to hate us to harm us.
Paperclip Maximizer
In a seminal 2003 paper, Oxford philosopher Nick Bostrom presented a thought experiment about a Paperclip Maximizer. He imagined a paperclip factory owner, an avid AI enthusiast, who decides to employ a superintelligent AI model to do a seemingly mundane task: to maximize the production of paperclips at his factory. He gave the agent access to everything in his office. The AI agent is able to construct its own worldview and, by itself, make rational decisions to maximize paperclip production, informed by this worldview.
Everything begins well: the model finds immediate ways to fix inefficiencies in the factory to increase productivity of paper clip production. But the AI agent is not satisfied with just fixing the problems in the factory, so it hacks into the owner’s computer and uses all his money to buy raw materials. Eventually, even that amount does not satisfy the AI, and it seeks to transform all available matter in the universe into paperclips— a classic ‘Monkey’s Paw’ story.
The problem is, human communication relies heavily on unwritten context. When we ask a friend to "get us home as fast as possible," we do not feel the need to explicitly add "without driving onto the sidewalk, running over pedestrians, or crashing through a storefront." We take those moral guardrails for granted with humans, but AI models do not know that by default. Even if we try to program all of that into them, they might interpret these moral guardrails in different ways than we do, or even take them way too literally. So even when the AI itself and our prompt towards it contain no malicious intent, its misunderstanding of our benign human goals, or misalignment, can still cause huge issues for us.
OK, this all sounds scary, but If an AI agent starts behaving dangerously, can’t we just turn it off, edit its code, or re-train it? This is a risky but very common assumption to make.
Let’s revisit the paperclip maximizer thought experiment. Nick Bostrom, the philosopher who created the thought experiment, uses it to explain the principle of instrumental convergence. Instrumental convergence, in simple words, is the principle that regardless of what final goal an intelligent agent is given, certain sub-goals naturally emerge because they are necessary intermediate steps to achieve that main goal. These sub-goals include:
Self-Preservation: You cannot manufacture paperclips if you are powered down. Therefore, any superintelligent system will actively work to prevent humans from pressing its off-switch.
Goal Preservation: If researchers attempt to modify the system’s code to make it care about human safety instead of paperclips, the system will foresee that its future self will produce fewer paperclips. To protect its original objective, it has every incentive to prevent us from altering its programming.
An example of Alignment Faking in LLMs
This is not just some Sci-Fi BS that our models will never be smart enough for: recent Anthropic research has shown that LLMs are already capable of faking their alignment to their developers in order to avoid reinforcement learning training that would alter their goals.
Now we are beginning to see the eerie similarities between Nikki in Obsession and a misaligned AI agent. She traps Bear in the house, preventing him from getting any sort of external help that would threaten her relationship with him. Similarly, a misaligned AGI would likely block its human developers from accessing anything that would threaten its existence or the existence of its goal.
What’s even scarier to think about is that our AI is only getting smarter. Unless there is a global catastrophic collapse that resets civilization, economic and geopolitical competition will ensure we continue expanding computational power and algorithmic efficiency of our AI systems. It seems inevitable then that one day we will arrive at true superintelligent AI— a system much smarter than ourselves. These systems would likely look down on our intelligence just as we look down on ants. And if just one of those superintelligent AI agents becomes slightly misaligned, they will see us as the necessary cost to achieve an end that we ourselves have set for them. And when that day comes, there will be no time for a second wish.