Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.


Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.


I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard.

With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or maybe just thinks for 10 minutes instead, then fixes one thing and breaks four additional ones.


yep matches my experience completely

But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.


Is there any model that knows how to smooth an overly literary text over? I find Opus and Fable constantly decorate the documentation they write like a damn 19/20th century writer. We're working with IT stuff yet it writes like it's going to win some Pulitzer prize. It's that one thing I don't get why they can't train them to do properly: I have not encountered a model yet that sticks to the current language of the domain it's tasked with.


Not sure how to fully fix this but I remember a session last week where I got so fed up mid way though reading a response that I used the following:

"give me this again without jargon invented this session at high density

and with a couple (maybe more or less) simple useful ascii diagrams underneath each design"

The context is that I was discussing an experimental new idea for my video game review analysis product.

Designs 1,2, and 3 were horrible: the model even suggested a rejection after the word soup so it would have been pointless to waste my fleeting time on Earth reading it.

Otherwise, I generally really enjoyed using fable for bouncing ideas. It was an absolute joy to have this thing provide useful criticism, analyse sample data, and create prototypes so that I could elevate my understanding of the problem without stepping down from a pure intuition/design headspace.

But I don't consider the purely model written code usable for a feature this important. I'll probably scrap it entirely and start from scratch with newfound understanding.


I've just seen it write this, I'm still laughing/crying:

> Monitor clipping. review_watched produces PathStatus and nothing else. It never feeds solve. Whatever it does to legs cannot reach the search.

(that's after being told twice to not use shorthand jargon nor reference the code directly)


Replying to myself, because I just bumped into these: https://www.reddit.com/r/claude/comments/1vfvdgz/anthropic_l... https://www.reddit.com/r/ClaudeAI/comments/1vgpyni/my_opus_5...

Especially the second one seems exactly like my experience.


The cursing thing blows my mind. "User is upset? Let's make decisions even faster (ie. more wrong) because clearly that's what they want!"

It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?


It's training data might have a ton of examples of people hurrying and screwing more after being yelled at


I've ditched Anthropic completely because of it. It makes me furious.


It's my daily driver. I like it and find it noticeably better than Opus 4.8.

After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.

My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.


> My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.

I’m doing the same right now, and I’ve found that asking for “simple English” works most of the times, although not always.

Did you find better wording that works consistently?


I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.


I have Fable plan and Opus implement. I haven't had any major issues working this way; however, Opus does seem plain fucking stupid compared to what I experienced with Sonnet previously.


I do the same, and generally have good results, but it does stupid things with gusto.

I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.


> however, Opus does seem plain fucking stupid

Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.


I bought my first LLM subscription with Claude right before they gave access to Fable 5.

I got so used to it, when they finally pulled access for me and I had to go back to Opus I felt like I was working with my hands tied.

I finally know what those women with AI boyfriends felt like when their app updated and it won't dirty talk with them anymore.


Fable has spoiled us all.


Whether this is true or not has been keeping me up at night the past week. Daily driving Fable is legit superpowers. Running out of usage and trying to work with literally any other model and everything breaks down because they can’t keep up without constantly tripping and derailing everything, meaning I’m working full time to babysit every judgement call they make instead of flying like a rocket.


Yeah, tell me about it. Anything else feels like what going to a local coding model used to feel like. Granted it still fucked up on occasion, but like maybe twice a week, not literally every other turn.


Not so sure, I'm sure Opus 5 is just shit.


Eh it's better than 4.8 in terms of what it can get done on a good day, it's just far more taxing to get it there.

Like the Fable ban stunt, I wouldn't put it pass Anthropic to kneecap Opus deliberately to drive more people to their more expensive option.


> I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff.

To me it's not so much the dumb mistakes (although there are some of those) but the ultra-verbose, mega-inefficient "solutions" to some problems / prompts.

Stuff that "works" if you're the kind of person that considers slamming a semi-trailer at 200 mph into a door did, technically, result in the door being somehow "open".

As it's supposed to be one of the most advanced model, I can't help but wonder if the solutions are that bad/verbose/inefficient because we're already in a loop of models being trained on sloppy-pasta from previous models.


Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.


Every model that comes out comes with a bunch of people saying "this one is actually dumb they were smart before" and I don't really get it. The models since Opus 4.5 have all been basically the same to me. Sometimes they do the wrong thing, so you have to steer and stop and correct them. Leaving them to operate on their own in no-human-in-the-loop harnesses often gets bad results. But if you single thread it, and keep your work targeted (you have to know what you want the thing to do!), clear your context, the models will do what you ask pretty reliably.


Glad to see this comment as this has generally been my experience as well. I'm really curious to see why it's so infuriating for others. My best guess is I'm using it more conservatively than most other users in this thread.


Statements like this typically come from working on the same setup and context using different models. I actually have that very experience now; I work on something security-adjacent so Fable often drops out, at which point Opus behaves like its lobotomized half-sibling. Pardon me the language, but I can't find a better example to be honest.


At this stage in the game almost none of the comments or articles on HN can be trusted, if you know what I mean...


Agreed. Opus 5 is doing just fine, slightly better than 4.8. It's personality is insufferable, but I find myself catching fewer problems at code review. It generally understands my conventions and isn't so eager to accrue tech debt.


Sent to solve one task, came back with half of it solved and 2 more problems.


What really enrages me is the amount of effort it puts into justifying weaseling out of work. (THAT'S MY JOB!)

It will do everything it can to defer or push it off, to the point where I’ve had to add multiple imperative directives to the AGENTS file telling it, in no uncertain terms, not to defer tasks under any circumstances.


sounds like someone needs a local llm.


Oh I do. Headless 128GB RAM machine serving llama.cpp with a number of local models that I use on a daily basis.

• Qwen3-VL picks up new images in a NAS, auto captions and adds the text descriptions as a hidden EXIF layer into the image, which is used for fast search and organization in conjunction with a Qdrant vector database.

• Gemma3:27b is used for personal translation work (mostly English and Chinese).

• Some small 8b models (like llama3.1) for sentiment analysis on text.

But haven't really tried using local LLMs in conjunction with agentic harnesses yet.


recommend opencode w/qwen 35B or 27B with MTP.

My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context. opencode's dynamic context pruning plugin can get you pretty far into the stratosphere.


Any experience with Sleev as replacement of DCP?

See https://news.ycombinator.com/item?id=48883538 25 days ago

> The Sleev (the project has been renamed to make a startup) creator was shilling their project in the OpenCode Discord. That person is very convinced they have something that no one has ever built before. They focused on token reduction without any real evals for capability impacts.

I'm generally against this context pruning without prompting or details. Sleev is very opaque about how it works and definitely will bust your cache.


No. the opencode plugin dynamic context pruning is essentially just labeling the context and then summarizing it; you can manually expand the context if you missed something, but most of the time, I'm using it to extend context into different scopes rather than trying to achieve the same goal.

There's a few times it gets too dumb to cope, but in comparison with other coding agents, it seems smarter.


> My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context

Thanks for the tip - I like this a lot. I remember having to do a lot of tweaking to curtail Qwen QwQ-32b when it would go down an endless psychotic recursive reasoning loops as part of its "chain of reasoning."


I tailored the agent to ao both its system prompt and budget-message align.

I havent yet tailored the pruning messages, but mostly it works.

Reasoning budget can also be set by client, so potentially smarter.


"One thing worth your attention", "Two things worth knowing", "One thing to eyeball"


And one of them is always something just completely out of scope and the other is something obvious it missed.

“One thing worth your attention, if you were to detonate a pipe bomb in your house, it would have a negative effect on your living room”.


This is driving me crazy. opus 4.8 did not do this to me not (at least during pre-5.0 timeframe). Feels like the new cycle is one step forward, two steps back.


I don't use Claude code, just Claude web, and I get this all the time. Or (since I have it push me to actually think) it will ask me some question in our back-and-forth, and then right after it'll provide the answer. As a "hint". Like come on


I have some internal tests I use for areas where one particular solution/paradigm is dominant but worse.

Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.

Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.

Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.


Yeah, I've dropped back to 4.8 entirely for the remainder of this billing cycle. I'm going to be seriously looking into Qwen adoption and harness migration options over the course of August.


Same.

I could not get Opus 5 to do anything without losing a few years of my life from stress.

Fable has been okay but I am doing ML work and not allowed to use it which feels insane.


Ironically, Opus 5 is the most benchmaxxed model I've seen from Anthropic. It is legitimately smart in a lot of ways but it has communication issues, both in terms of how it communicates (all the autism of GPT class models, without the brevity) and how well it catches all the nuance of what you tell it.


"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."


This is the hard-won load-bearing quote.


Belt and braces all the way down


The shape of this problem is very heavy


Yeah, I told it to save in its memory that I don't want to have any more word salad!


https://claude.com/blog/the-new-rules-of-context-engineering...

I don’t see anyone talking about how you have to completely change your prompting strategies with Op. 5 versus 4.8 to get the most success.


It works fine for me. Only issue I have is that it has me constantly reaching for the dictionary.


It is infuriating to interact with, but it is also first in many blind test leaderboards on LLMArena


What's the clear best, that you see?


What I’ve seen from LLMs is you have no idea which model does the best at a certain task until you try all the models. Gemini and Chat seem to be the best research models but also hallucinate like crazy. Oddly, Deepseek v4 Flash is the best web research model I’ve used. No idea why! There doesn’t seem to be a “best at everything “ and you don’t know which is the best at the thing you are doing until you try doing it.

Note this is for complicated tasks - simple stuff can be done by whatever pretty well.

Also models will be awesome at doing something at 150k tok of context and terrible at 750k toks so even within the model itself there are capability considerations.

It’s an engineering problem to design around, not something you can escape.


Hilarious to see only different responses


Should have been “clear best and what do you do”


GPT 5.6 Sol


GLM 5.2


Fable 5


Can not confirm, for me it's the complete opposite.


It's an infuriating model




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: