At least with Codex, this has not been my experience at all. It still screws up sure, but in every case I can ask "why did you do this" and it can trace back what made it take that particular decision. Typically it's always that I either didn't specify the problem correctly or made a really dumb mistake (executing the task on the wrong project....did this one yesterday) or it's something within a skill file that instructs it (at which point I fixup the instructions).
Once in a blue moon it's actually the model making a material error in it's thinking and I have to go back and redo it.
agreed to some extent. I think this parody still highlights what I feel is often the experience. It might not happen on a simple task such as changing a button color, but on more complicated things, this can definitely be exactly what it feels like.
Claude is great, but I have come to really hate the way it "talks". It's so irritating and there seems to be no way to make it speak normal English. So many claudisms in every response
It's a joke like the endless conservative dudes doing the "ordering coffee" joke is. It relies upon the ignorance of the viewer -- which is usually a fair assumption -- and basically that your understanding of something is based upon the prior accrued layers of "jokes".
Haha this is spot on how I've been feeling lately. I find it unbearable to work with this model for this reason... any trick out there you can do to steer it not to overcomplicate things? I guess Codex here I come
This got me on "cyanide blue", and I was ROLLING ON THE FLOOR LAUGHING on "Approaching usage limit". I can barely stop laughing now and my stomach hurts. I mean, Thank You!
To be fair I’ve worked on human programmed systems where similar “it should be a half point story” requests would be met with snark by the engineers and take 2 sprints.
I only made it through the first round of prompt selection; both options for the second step were equally pointless and not at all prompts I would ever expect to result in a constructive outcome. In my experience, telling the model it screwed up without specifically addressing, unambiguously, how to fix it, only leads to more suffering. If this page illustrates nothing else, I think it shows the immense downside of trying to use simple one or two sentence prompts.
So this site is just a fan-fiction that thinks it's somehow dunking on Claude? I've never had a session that remotely resembles any of this. I honestly can't tell what point this site thinks it's making.
I wrote my own harness to stop shit like this from getting to my attention out of frustration.
I'm sure there's quite a bit of variation from person to person in these sorts of experiences, based on your harness, the way you talk, the stored memory, your CLAUDE.md, etc. But people absolutely have had this Opus 5 style experience the app simulates.
Ok, I can't describe this, but my toes curled at the prompts, I was literally yelling at the screen. "None of these prompt options do what I want, they're going to make it WORSE!"
And yes, yes they did. The prompts options aren't ... diagnostic shaped? Some seem like they'd actually push the agent away from the solution.
One day we'll have proper theory of this and we'll laugh at people putting in magical input that obviously can't work.
Alternately models will get smarter and can read the problem space better despite the prompt, instead of because of it.
Yeah, I don't mind using AI to help me at work, but having to "talk" with this stupid crap all day will send me to an early pension or something. Can't be healthy in the long run.
Exactly! Never thought of it, but now that you say it. You know, like when you talk to a person in a neutral environment and you just can tell that they are a school professor for example, because of the way they express themselves. Now I wonder what imprint Claude leaves on us.
P.S. I have utmost respect to school professors by the way. My mum is a prof and I see it firsthand in many of her colleagues too. I guess explaining same things multiple times a day and talking to kids only for 8 hours straight objectively creates some recognizable speech patters and habits.
I’ve never used any of these tools. Please tell me that this is a grossly exaggerated parody, and that the tools don’t write like this, or do so many ridiculous things. For my sanity.
(I am genuinely uncertain, though I presume it’s at least somewhat exaggerated.)
I use Opus and Sonnet 5 all the time and I find their language grating. But honest, I prefer to put up with it and get the results than to put up with my own human limitations and not get the results.
I heard this was a thing when listening to a Theo podcast, he mentioned to add a "Only use subagents if the user explicitly requests them" line in your agents.md file.
I don't know if it works, but I've always had a consistent level of token burn on my plans (I've only heavily used Sol after adding it).
I dunno, in my experience it's often overly-narrow, sometimes jumping through all kinds of hoops to preserve some edge-case behaviour that doesn't matter because I didn't mention it could be changed.
I was expecting it to spend 30 minutes running headless chrome instances, taking screenshots and analyzing them in python to verify the blueness of the result.
Nails the Claude dialect. Technical nonsense like:
> I'm collapsing this back to the rendered outcome:
And intermixed with SaaS product page idioms from a brain-damaged marketer like:
> No broader cleanup.
> No further architecture work.
> Just the button.
Aside from the patterns everyone knows like em-dashes, "its not X, it's Y", etc. I think the key features of claude diction is it sounds like a junior engineer over their skies who is trying to make up for that with extra verbiage mixed with extremely grating SaaS marketing-ese.
Am I the only one whose experience doesn't match this?
My gripe with Claude is that while investigating how to do this it will report 200 other incidental findings which I overlooked and I realize those are broken too and need urgent fixing, derailing me, not it.
Oh man, exactly. I'm very prone to scope creep as I work on tasks. I already would notice some things that could be fixed or refactored and have a hard time not touching them before I used agents. But now I have to be very intentional about not letting it manipulate me into fixing EVERYTHING RIGHT NOW. Half the time the "one more thing worth noting, unrelated..." isn't even an actual issue, it just brought it up to fish more usage out of me.
Also, while this little demo is certainly exaggerating the issue, I do find working with Claude to sometimes get quite verbose and tiresome. I doubt I would struggle this much to get it to change a button color, but the patterns of speech, the endless lists, the over-explanations, and the whole song and dance of trying to get it to make the change you want without side-effects is frustratingly familiar to me.
Just when I came back to my pc and was thinking "I hate this world were everyone talks about AI like fanatics" this made me a little bit happy, especially the unhingend all caps options towards the end
At least with Codex, this has not been my experience at all. It still screws up sure, but in every case I can ask "why did you do this" and it can trace back what made it take that particular decision. Typically it's always that I either didn't specify the problem correctly or made a really dumb mistake (executing the task on the wrong project....did this one yesterday) or it's something within a skill file that instructs it (at which point I fixup the instructions).
Once in a blue moon it's actually the model making a material error in it's thinking and I have to go back and redo it.
agreed to some extent. I think this parody still highlights what I feel is often the experience. It might not happen on a simple task such as changing a button color, but on more complicated things, this can definitely be exactly what it feels like.
I got way too annoyed at this before realising it was an optional game and I could just close the tab
That's so weird... This doesn't at all match my experience with Claude. I've never seen it behave this way.
You were right to push back. It's not an accurate representation of claude. It's satire.
Seriously, I use Claude Code all day and have zero issues with this.
To the people rather lamely doing the "it's satire/a joke", that would require this to be an exaggeration of a reality. But...it isn't.
Claude is great, but I have come to really hate the way it "talks". It's so irritating and there seems to be no way to make it speak normal English. So many claudisms in every response
try using Claude Design
It means that either you stopped using Claude around Opus 4.6 or you use Fable instead of Opus 5 :)
I use Opus 5 for everything.
because you never changed just one button to blue
It's a joke dude.
It's an insult to the superintelligence. The basilisk will not look kindly on this!
It's a joke like the endless conservative dudes doing the "ordering coffee" joke is. It relies upon the ignorance of the viewer -- which is usually a fair assumption -- and basically that your understanding of something is based upon the prior accrued layers of "jokes".
"it's funny because it isn't true"
Yeah, but it's not funny since it doesn't match reality.
This is actually what keeps people using AI: variable reward schedule. It's basically gambling.
People say this, but I've never seen it. AI has been very consistent in its rewards for me.
Which also explains why response speed is so important.
Haha this is spot on how I've been feeling lately. I find it unbearable to work with this model for this reason... any trick out there you can do to steer it not to overcomplicate things? I guess Codex here I come
`/model claude-opus-4-7`
-= CAUTION, SPOILERS =-
This got me on "cyanide blue", and I was ROLLING ON THE FLOOR LAUGHING on "Approaching usage limit". I can barely stop laughing now and my stomach hurts. I mean, Thank You!
To be fair I’ve worked on human programmed systems where similar “it should be a half point story” requests would be met with snark by the engineers and take 2 sprints.
I guess we are all PMs now.
The user prompts are as comically bad as the responses, the classic slapstick dynamic I guess
I don't get who this is making fun of:
- The people who won't make any effort to learn the tools, and something as simple as reverting code (via git) needs to be done by AI?
- The awful programmers who we've had to endure working with, who are so bad at simple changes that they have negative productivity?
- Or Claude itself?
---
BTW: I don't have these problems, but I'm also not afraid to do things myself when it's easier.
I lost it when it finally did the right thing, but then it added a never-requested gradient to the button. Very good!
I’m impressed you had the patience to even make it that far!
I only made it through the first round of prompt selection; both options for the second step were equally pointless and not at all prompts I would ever expect to result in a constructive outcome. In my experience, telling the model it screwed up without specifically addressing, unambiguously, how to fix it, only leads to more suffering. If this page illustrates nothing else, I think it shows the immense downside of trying to use simple one or two sentence prompts.
I never really understood what being "triggered" was like until now.
I was waiting for it to .. say usage limit reached after reverting it back to how you started..
Was this made by someone who hasn't actually used any of these tools in over a year?
So this site is just a fan-fiction that thinks it's somehow dunking on Claude? I've never had a session that remotely resembles any of this. I honestly can't tell what point this site thinks it's making.
So you are lucky, congratulations.
I wrote my own harness to stop shit like this from getting to my attention out of frustration.
I'm sure there's quite a bit of variation from person to person in these sorts of experiences, based on your harness, the way you talk, the stored memory, your CLAUDE.md, etc. But people absolutely have had this Opus 5 style experience the app simulates.
Ok, I can't describe this, but my toes curled at the prompts, I was literally yelling at the screen. "None of these prompt options do what I want, they're going to make it WORSE!"
And yes, yes they did. The prompts options aren't ... diagnostic shaped? Some seem like they'd actually push the agent away from the solution.
One day we'll have proper theory of this and we'll laugh at people putting in magical input that obviously can't work.
Alternately models will get smarter and can read the problem space better despite the prompt, instead of because of it.
but still.
I’m laughing and crying at the same time. This is what work feels like now. Thank you, well done!
Yeah, I don't mind using AI to help me at work, but having to "talk" with this stupid crap all day will send me to an early pension or something. Can't be healthy in the long run.
Exactly! Never thought of it, but now that you say it. You know, like when you talk to a person in a neutral environment and you just can tell that they are a school professor for example, because of the way they express themselves. Now I wonder what imprint Claude leaves on us.
P.S. I have utmost respect to school professors by the way. My mum is a prof and I see it firsthand in many of her colleagues too. I guess explaining same things multiple times a day and talking to kids only for 8 hours straight objectively creates some recognizable speech patters and habits.
I’ve never used any of these tools. Please tell me that this is a grossly exaggerated parody, and that the tools don’t write like this, or do so many ridiculous things. For my sanity.
(I am genuinely uncertain, though I presume it’s at least somewhat exaggerated.)
No. It’s just a bunch of jokes rolled up into a big exaggeration.
It’s funny because there are elements of truth in each bit of it, though.
Brilliant. Precisely the reason I stopped using Anthropic's products.
That was funny :-)
I use Opus and Sonnet 5 all the time and I find their language grating. But honest, I prefer to put up with it and get the results than to put up with my own human limitations and not get the results.
> 23 agents total.
This hit a bit too close to home. Sol has the same issue, spawns a lot of agents for no good reasons (besides burning tokens).
I heard this was a thing when listening to a Theo podcast, he mentioned to add a "Only use subagents if the user explicitly requests them" line in your agents.md file.
I don't know if it works, but I've always had a consistent level of token burn on my plans (I've only heavily used Sol after adding it).
This felt very late-90s net art. Stressful but nicely done satire.
Congratulations!!! You win what’s left of the internet - just ask Claude for your prize! Motrin I’ve had this week.
I use codex now.
This does not match reality at all, speaking as the #1 user on agent hours per clauderank.com
I find it funny how it went off with subagents and adversarial review when a simple grep or diff is sufficient.
It's funny but unrealistic as Claude does a pretty good job at only changing what is required these days with the 5 tier models like Opus 5 or Fable.
The site is opusfived.com, and Opus 5 is probably the worst so far at doing this.
I dunno, in my experience it's often overly-narrow, sometimes jumping through all kinds of hoops to preserve some edge-case behaviour that doesn't matter because I didn't mention it could be changed.
Experiences vary, yes.
But you can see in this thread that folks definitely have experienced this.
This is exactly the experience I have with Opus 5. Opus 4.6 is better, Flable 5.1 much better. But Opus 5 is infuriating.
I was expecting it to spend 30 minutes running headless chrome instances, taking screenshots and analyzing them in python to verify the blueness of the result.
How do you manage your frustration in these interactions? I often find myself getting pissed off
I stop using AI and do the job manually. I normally give AI one shot at the task. If it fails then it's not saving me any time
The goal of my personal harness is to get to the point where I never actually talk to Claude directly for that very reason.
I don't get the joke... maybe because I'm using Codex?
its making both buttons blue
This is so good at replicating the experience of frustration, then relief when it finally does what you asked it to do in the first place!
This is so perfect and depressing that I might cry. It’s like a Kafka novel about programming.
I've experienced this so many times over.
"I was wrong" and "the honest truth" are just forever phrases that are now dead to me.
I want the dishonest truth.
Username may check out
it gave me headache in 2 turns, just like opus 5 !
This is scary close to my interaction with Claude this week.
glad to see I'm not the only one... anthropic needs to support my anger management treatment
This is pure genius. No notes
This spiked my blood pressure. Well done
Nails the Claude dialect. Technical nonsense like:
> I'm collapsing this back to the rendered outcome:
And intermixed with SaaS product page idioms from a brain-damaged marketer like:
> No broader cleanup.
> No further architecture work.
> Just the button.
Aside from the patterns everyone knows like em-dashes, "its not X, it's Y", etc. I think the key features of claude diction is it sounds like a junior engineer over their skies who is trying to make up for that with extra verbiage mixed with extremely grating SaaS marketing-ese.
This is gold, thanks for the giggles! I think it was designed that way to burn tokens.
Here come all the totally organic "wow, I guess I better switch to OpenAI" comments.
Am I the only one whose experience doesn't match this?
My gripe with Claude is that while investigating how to do this it will report 200 other incidental findings which I overlooked and I realize those are broken too and need urgent fixing, derailing me, not it.
Oh man, exactly. I'm very prone to scope creep as I work on tasks. I already would notice some things that could be fixed or refactored and have a hard time not touching them before I used agents. But now I have to be very intentional about not letting it manipulate me into fixing EVERYTHING RIGHT NOW. Half the time the "one more thing worth noting, unrelated..." isn't even an actual issue, it just brought it up to fish more usage out of me.
Also, while this little demo is certainly exaggerating the issue, I do find working with Claude to sometimes get quite verbose and tiresome. I doubt I would struggle this much to get it to change a button color, but the patterns of speech, the endless lists, the over-explanations, and the whole song and dance of trying to get it to make the change you want without side-effects is frustratingly familiar to me.
Claude is an unbelievable yak shaver if you let it be.
You're not alone. I've been sat wondering what kind of codebase someone has if they have this problem, I've never seen this behaviour.
Just when I came back to my pc and was thinking "I hate this world were everyone talks about AI like fanatics" this made me a little bit happy, especially the unhingend all caps options towards the end
PTSD 9000.. I miss the old days, less load bearing BS and more in the zone coding..
Fair play. There's a quiet truth to what you're saying, and it's worth pointing out
now THAT is a load-bearing simulation
Lol this is great
rofl, brilliant