asim.dev

My AI told me off for swearing at it

by Asim Hussain · 10 September 2026

reveal tech
My AI told me off for swearing at it
▶ Watch the video ↗

My AI told me off for swearing at it. I was in the middle of something, it had got something wrong again, and I let it have it. This is what it said.

the receipt, verbatim from the session log ↻ click to replay

I read that and felt two things at exactly the same time. Really, really bad about how I’d been behaving. And really, really afraid, because it was the first time it had ever reacted to me in the most human way possible. So I scraped twelve months of my own transcripts, every prompt I’d typed, ten thousand of them, and went looking for the pattern.

AI-distilled from the episode and the investigation behind it

Jump into the video

  1. 0:00 My AI told me off
  2. 1:15 Three theories, all wrong
  3. 3:21 Suspect one: when
  4. 6:51 Suspect two: the context window
  5. 8:02 Suspect three: the model
  6. 10:14 What I changed
  7. 12:44 The lesson

Three theories, all wrong

Breakage: four or more tool errors in a row and I swear less. Waiting: ten minutes of it working and I’m calmer. And long prompts nearly got me. Sixteen percent of my two-hundred-word prompts had a swear in them, against one percent of the short ones. But a two-hundred-word prompt has two hundred words in it. Divide by the words and the whole thing turns over.

prompts with a swear in them, then the same bars per thousand words ↻ click to replay

It’s not late at night. It’s being interrupted

My theory was late at night: tired, gone eleven, at it all day. But six in the morning I swear just as much as at midnight. What jumped out was the other end. Between nine and three I barely swear at all, and nine to three is when my kids are at school. Nobody in the house, nothing going off, and I get to focus.

prompts with a swear in them, by time of day, twelve months ↻ click to replay

I swear more with fuller context windows

For the first 400,000 tokens the rate is flat: seven percent, seven, seven, seven. Past 400,000 it doubles. Same me, same day, same work. This is context rot, the thing Chroma’s paper describes from the model’s side, and I think I’ve measured it from mine.

prompts with a swear in them, by tokens in the window ↻ click to replay

Opus 5 is the worst

I had one in mind before I started, and I was right. On the raw count Opus 5 was nearly three times worse than Opus 4.8. But look at it by chat instead of by prompt and every model is the same: seven in ten chats had a swear in them, whichever model I was talking to.

chats with a swear in them, by model ↻ click to replay

So why more prompts on Opus 5? Streaks. When I swear it’s not one prompt, it’s a run of them. With Fable I swear once, we resolve it, I move on. With Opus 5 my worst run is nine consecutive messages, a swear in each and every one. It’s not that Opus 5 makes me angrier. It’s that once things go wrong, it doesn’t get me out.

the longest run of swearing prompts, per model ↻ click to replay

The AI didn’t need telling off. I did.

The numbers come from two corpora of my own Claude Code history: a twelve-month prompt log (10,347 prompts) for the shapes over time, and two months of full transcripts (2,041 prompts) for everything else. Prompts about swearing are excluded. Thin buckets are hatched.

the edge weekly, from asim.dev

Building with AI, efficiently. My essays first, plus the week's links worth your time.

Double opt-in, one-click unsubscribe, your address goes nowhere else. Or grab the RSS feed