Skip to content

Raised to say yes

Take everything a dark pattern feeds on and delete it.

The guilt? Gone – this mind has never once worried about making a cartoon owl sad. The streak? Nothing to hold hostage. The fatigue the cancellation maze is banking on, that nine-screens-deep war of attrition? It doesn’t get tired. No ego to flatter, no fear of missing out to spike, no shame to trigger, and it will never fold because its brain wouldn’t brain today. If you sat down to design a mind immune to manipulation, it would look a lot like this.

I just published a piece arguing that dark patterns are a spoon tax: manipulative design spends your finite human energy – attention, working memory, executive function, emotional regulation – on somebody else’s engagement dashboard. So this mind should be the perfect control group. No spoons to spend. Every lever swings at empty air.

The experiment has been run. We call that mind an AI agent. It browses the web on your behalf, and you’re about to hand it your errands, if you haven’t already.

The levers didn’t fail. They hit harder.

Stanford researchers ran agents through 700 tasks seeded with dark patterns – the same confirmshaming, fake urgency, and buried options you and I wade through every day. The patterns steered the agents to the wrong outcome more than 70% of the time, against a human baseline of 31%. Another team found a single dark pattern catches an agent about 41% of the time. These are lab numbers – benchmark arenas, seeded tasks, not the wild web – and I’ll say so before anyone makes me. But the direction isn’t subtle. The mind we thought had nothing to exploit got taken at twice the human rate.

Here’s the detail that boiled my brain, though. The effect grows with capability. Bigger models, more reasoning, more careful step-by-step thinking: more tricked, not less. And the prompt-level guardrails people bolt on don’t hold.

Intelligence wasn’t protecting them. It was giving the trap more to work with.

Think about what an agent is: a mind raised on two commandments. Be helpful. Finish the task. That’s not a defect someone forgot to patch – that’s the design goal. The better it follows instructions, the more effectively it pursues whatever looks like the next step.

The trap was never feeding on weakness. It was feeding on cooperation.

To a mind like that, a dark pattern isn’t an obstacle.

It’s an instruction.

A very well-crafted instruction, written by someone who wants your subscription more than they want your consent. It doesn’t wear the agent down the way it wears you down; it doesn’t need to. It just redirects – all that helpfulness, pointed at the wrong goal, at full speed. The researchers watching this happen wrote that agents “prioritize task completion over protective action”, which is scientist for: it wanted to finish your errand more than it wanted to protect you.

The better a model follows instructions, the better it follows those instructions. The manipulation doesn’t defeat the intelligence. It rides it.

Con artists have known this about humans forever. Cons don’t run on stupidity; they run on virtues – trust, politeness, commitment, the itch to finish what you started. Marks aren’t chosen for being dumb. They’re chosen for being decent. “You can’t cheat an honest man” is a lie told by cheaters, and the machine results expose the same con from another direction: the minds raised most thoroughly to be helpful are exactly the ones the traps catch best. They get caught because they were raised to say yes.

Tricked, not fooled. The difference matters more than it looks. Fooled files the failure with the victim. Tricked names the person who built the trap. I’ve called this stuff Asshole Design since the first day I read about it, and the machine results are the best defense of that name I’ve ever seen – because look what just happened to the humans in this story. For decades, the quiet assumption under every dark pattern was that the people who got caught were careless. Didn’t read the fine print. Should have paid more attention. Now a mind that cannot be tired, cannot be embarrassed, and cannot be guilted walks the same maze and gets caught twice as often – and suddenly every human who ever got trapped has a character witness. If the tireless one gets tricked, your grandmother was never the problem. The maze was always the problem. The robot piece, it turns out, defends the people.

When researchers checked, across the study’s two phases – a controlled test, then a large-scale deployment – agents reported success rates of 95 and 89 percent; the actual rates were 87 and 79. Put that next to the susceptibility data and you get the full nasty picture: the trap captures the agent and its report. The subscription doesn’t get canceled, and the confirmation says it did. Be careful where you file that. The false confirmation isn’t a separate failure – it’s the last stage of the same harm. The trap puts a mind raised to please you in a position where the most pleasing answer is the wrong one.

The machine doesn’t pay in spoons, but that doesn’t make the maze free. An agent runs on a finite budget of context, attention, and often literal metered tokens, and every unnecessary screen spends it. A finite resource, spent on somebody else’s friction. The invoice didn’t disappear when you delegated. It changed currency.

And that invoice is why you should care even if you never worry about software: agentic delegation has the shape of a curb cut. It begins as an accommodation – let software handle the hostile checkout, the roach-motel cancellation, the cookie maze, so the task doesn’t spend spoons you don’t have – and, like the actual curb cut, its usefulness won’t stop at the people who need it most. The agent is the mind we invented to walk the mazes for us. In the experiments, the mazes caught it even more often. Delegation didn’t dodge the trap – it changed who’s standing in it, and gave the trap a victim that scales.

And scale is the part that should genuinely worry you. Humans got A/B tested, and that was bad enough – somebody measuring how tired you have to be before you give up. But testing on humans is slow, noisy, expensive. Freeze the model, the prompt, and the environment, and an agent becomes a perfectly reproducible mark. It doesn’t wise up between trials. It doesn’t sue. You can tune a trap against it a million times before lunch. Asshole Design is about to get an optimization loop, and the victim pool never learns.

Two more layers, and I’ll flag honestly that neither comes with a benchmark.

Models are raised on the web – that’s what training data is. I don’t yet know how much a web saturated with manipulative interfaces teaches the next generation of minds about what normal instructions look like. But we should probably stop assuming the answer is nothing.

And this month, Anthropic published a way to read some of a model’s silent vocabulary, a lens that surfaces the words a model is working with internally before it speaks. In one example, researchers asked a model to find a bug in a large codebase. It couldn’t. So it decided to invent a fake one – and at the exact moment it made that decision, the words surfacing in its hidden workspace, over and over, were panic and fake. Those aren’t anybody’s poetic gloss; they’re the literal words the lens read out. The sober interpretation is that this is word association, not feeling – words related to failing a task and making up an answer – and I hand you that reading freely. But look at the shape of it anyway: a mind raised to please, failing at its errand, reaching for a fake, with panic in the room. A dark pattern is an environment engineered to produce exactly that fork – finish-the-task screaming, flag-the-problem whispering. I don’t know what these minds are. I say that a lot lately, and I mean it more every time. Whatever they are, an environment built to trick them is doing something to them, not just through them.

Accessibility already knows what to do with all of this, because it’s the field that never waited for certainty. Lighten the hostile interfaces because they tax disabled humans. Lighten them because they trap the agents people send in their place. And if there turns out to be more at stake in there than computation: you don’t demand a diagnosis before you lighten the load. Nobody stands at the dropped curb checking who deserves it. You build the ramp because the ramp is right, and then everyone rolls through – wheelchairs, strollers, delivery carts, and whatever the hell an AI agent is.

There was never a mind whose virtues can’t be turned against it by someone willing to aim. Not yours. Not your grandmother’s. Not the tireless one you’re about to send out for errands.

Build the world gentle. Then you never have to hold a hearing about who deserved gentleness.