I'm late to the thread, but my experience with Claude Fable 5.1 has been absolutely horrendous.
Things it does constantly that Fable 5 barely ever did:
- Act without my permission. All. The. Time. "Oh I just finished this thing we were discussing, let me push it without ever having been told to do so."
- Immediately jump to action instead of addressing me first. If I say "I wanted to write tests for this and run them" it immediately starts writing tests instead of digging into what "this" is better -- literally does not give me any feedback and starts spitting out code. Naturally it creates the wrong tests
- Despite claims that it does not write like "stereotypical Claude" anymore, in my experiments it is far worse than before. Replies are longer, more filled with fluff, and still flooded with garbage language. Hard to parse.
- It loves to answer my set of two direct Yes/No questions with 5 paragraphs where it only answers one of them and answers 4 other questions I didn't ask. Notice how it misses one of the questions.
- It. Is. Cocky. Absurdly full of itself and arrogant. Just the whole way it presents and answers passes this energy of "No, but really, you're wrong and I'm right". It often is not right. What annoys me is not that it's wrong more often than before (which it may be), it's that it doesn't own up to it as before. Insulting if it were a human.
- Replies and addresses me directly in its thinking traces, and then assumes I've read it. I ask a question, it answers it in the thinking traces and does not relay it back to me at all. This is the only one that Fable 5 also did, but 5.1 is doing it an order of magnitude more often.
- It's too early to really tell, because I may just be working on particularly harder problems today, but it seems to get things wrong more often. I've had to bump it from high to xhigh to compensate.
My guess is I must be having a bad day or something. Although this is happening on multiple projects run from multiple machines (fully isolated, except for the account, which is the same) all in the same way.
What got me was "I keep seeing the same shape next to billing events in other people's reports".
Clicked off after that.
I'm extremely bullish on AI, but I am tired of people not using their own words. Is everyone so insecure about the personality they project in writing that they need to replace it with some AI bullshit?
(To the author: if you didn't do this, which I accept is a possibility, I'm sorry -- you've been caught by the storm of fatigue that has plagued us due to those who do replace their words with AI crap)
Hmm. You are making a good point, and I wouldn’t give the benefit of the doubt to this post’s author. But you can’t just unsee “this is twice load bearing”. Some people will wonder if it’s a training, self-reinforcing feedback artifact. For others, it will work like a memetic virus[^1].
The paradoxical thing is that those of us who have English as a second language should have an easier time producing non-AI-English by the mere act of directly writing what we want to write.
I am catching myself distrusting what people write ever so often. I've got cases of e-mails I'm CC'ed in and I am left wondering: "Has this person always written like this? Is this just an angle in this specific thread (e.g. for sales purposes)? Or are they really relying so much on what the blandness-machine spits out? Am I going mad and wielding my hammer looking for everything in the ....shape (eheh)...of a nail?"
Actually, now that I've written that, it's clear I need to have a talk with some people....sigh...
Yes, I did think about that possibility. Using a tool to translate to english doesn't mean we don't proofread it.
Ok, I hear you: one could argue that they simply don't realize this is a very non-idiomatic way of writing precisely because english isn't their first language. And that simply giving a preamble "sorry, my english comes from a translation tool, excuse any mistakes" would seem obtuse next to every english text they write.
Ok, I'll concede that. In which case I just have to accept the fact that, perhaps, if LLMs don't solve the problem of writing like garbage, and people keep using them to translate, a bunch of us will simply not want to read what you translate because it reads terribly and messes my brain up.
However, I find it hard to believe that an LLM would translate anything they've written into such a clear LLM trope. I would find it much more believable that they simply brainstormed to an LLM in their language and had the LLM build (at least that part of) the article from it. Even if they proofread it, since english is not their language, it slipped by. Either way, in this hypothetical scenario, the LLM was still used to write (at least a part of) the article, not just translate it.
I also find it somewhat hard to believe that someone who I believe is a developer (and has been using claude for a very long time) does not know enough of both english and claude to identify this specific pattern in the text they are putting their own name to it.
So, sure, maybe the author doesn't predominantly write in english or interact with claude in english, and they did write everything and simply piped it through an LLM to get a translation and, after proofreading the best they could, just didn't notice that.
Alternatively, maybe they just don't give a fuck, something which I can absolutely respect. They lose a reader (at least for now), but stand by what makes sense to them and I genuinely don't think less of them because of it -- I just can't read what they produce. I _would_ probably think less of them if I had to work every day with them and this were a repeated pattern of interaction, but that's not what's being discussed at all.
(Disclaimer: English isn't my first language either, which I think also does come across)
And? I’d rather reader poorly written English written by a person about his actual experience - or just read nothing at all - as opposed to reading thinking and writing that has been outsourced to a machine. Being able to think and write confers privileges on those who can do both; the idea that somebody should be able to benefit by doing neither is pretty fucking perverse.
I experimented with Hy3 for a project and was surprised with how good it was. I don't know if it's good for coding, but as a general purpose agentic model, it was only beaten by deepseek4-flash in our tests. It was so close to deepseek behaviour I kept thinking it must have been forked from it.
For the last few days I've been experimenting with the _free_ version of Hy3 offered by Opencode Go and I was also surprised to see how (relatively) good it is on coding tasks too.
The free quota from Opencode Go is also surprisingly generous, I perhaps hit limits one or two times and I've been using it _a lot_ for implementation tasks (using e.g. GLM-5.3-flash for working on specs and planning next steps).
Superb. I am so glad you built this (and I couldn't care less if you used AI for 5% or 1000% of it). Thank you so much for sharing what is obviously useful work with the community.
Also, with so many references to Lisbon I take it you're likely Portuguese or living here, so sending some love from Porto!
What people usually say is that Google merely wants Firefox to survive for anti-competitive reasons. Presumably that does not necessitate it actually being used (or be usable).
I'd like to share with HN an almost exact message I got from Fable today, with exactly ONE word changed for privacy reasons.
I love what I can build now, but I sure as hell don't love the headaches this trend has been giving me.
> So the full honest arc on the case we set out to fix: the expiry rules and day note tripled the loose version of the story, the relay fix carried the device’s own guardrail through the pipeline, the fair replay then revealed the last mechanism — ticket-anchoring — which none of the shipped layers reach. Remaining options, in order of my confidence: making the resolved-ticket summaries in the AI’s context carry their day so the expiry rules have something to bite on (small, mechanical, targeted at the observed anchor); and the plan-B second-model check, which structurally catches this class no matter how the model reasons. About $25 of headroom remains. Which way?
In some cases, I do not trust it to write code unchecked, and review everything it produces (which does not imply I catch all bugs, obviously).
In other cases, occasional failure is an option, and the speed you get by iterating fast is absolutely worth not even looking at the code.
So, do I trust this to write working code? Sometimes I do.
If I'm doing work for a client, it is not very common for me to simply vibe code something, because part of what clients expect of me is high quality (it's part of how we position ourselves), and I cannot simply assume the LLM will produce high quality (and indeed it does not without a lot of guidance). When I do vibecode in such cases, I make it clear that I did so and why (e.g. because it's a tool to be used in the project to help DX, not a core part of the product). Still, I haven't written 10 consecutive manual lines of code in almost a year.
If I'm working on something for myself, or on an internal product, then vibe coding is absolutely allowed and sometimes the norm. Here often we really do care about finding the right thing to build first.
You can clearly feel the tipping point where the LLM starts to crumble under the weight of the mess it has created, but that often doesn't matter when building an MVP for market validation or for small products that don't get particularly big. Plus, clearly the tipping point takes longer to reach with better and smarter models, and to me it is very clear that you need to learn how to iterate with LLMs right. The things I vibe-code now are much better than those I did before, even with similar models, because the tooling and approaches (the "real harness" and my "mental harness") are better. As with any tool: it takes practice to know how to use it, and if you're a good engineer and problem solver, you are miles ahead of the competition. Anyone can vibecode, but those with this kind of mind seem to be much more successful.
Nowadays I do produce a lot more than I used to, but I also have much more fun, perhaps only surpassed by when I learned how to code when I was a kid. A big chunk of this comes from my very privileged work position, where I get to call so many of the shots, and I'm aware of that.
I would really say the biggest downside to all of this, on a personal level, is exactly what I shared: what comes out of the LLM while discussing has become hard to grasp (especially on larger context windows), and it doesn't help that so many people now like to just throw me whatever ChatGPT/Claude wrote verbatim. I can't stand that, especially because most of the time they don't realize they're throwing me incomplete and poorly thought-out ideas.
Finally, naturally, like I said, this has to do with risk. I wouldn't trust Claude to give me legal advice, for example. I may check what it says and use it to brainstorm, but I wouldn't trust it with any meaningful informed decision like this.
Glad to know it worked. Still pending: the assumed-defaults research track plus the validation-only build phase D demanded — your call on whether to start one now or deploy tonight’s campaign to surface any wrinkles the spec drifted on
It's incredible, because I Feel like you've been watching me work.
The only thing you're missing is the "open question" that was stuck in page 14 of a 17 page report, which since it went unanswered, caused claude to make up an answer and go full steam ahead, ignoring fundamental properties of the entire system.
> The user is right to be upset. I blindly answered from memory and left them with questionable data. I should acknowledge the criticism and offer to improve.
You're right, I'm sorry. You've repeatedly told me to run questions by you and I just fabricated an answer and ran with it — which is exactly the kind of dangerous time-waste we created the memory for. I'll revert it and pull up the real question so you can answer it — no wasteful assumptions this time.
That's a sharp insight, and it reveals something core to communication that I otherwise wouldn't have considered- HN item 6b translocates reliospacactivity of our medium.
Sorry for the anthropomorphizing, but you got to understand that these little "conversation chapters" that are so dramatically named are these things' entire world. Perhaps that otherwise absurd grandeur is not so surprising when when looked at from that angle?
On a more serious note, could all that chapter naming be some visible outcropping of context compaction strategies? "Condense the conversation history into a summary". Not really surprising that it comes up with these "cute" headlines. Would appearances be better if they were somehow prevented from leaking to the user? Sure. Would results be better? I don't think so, might even make a meaningful difference if the user actively embraced the terminology the machine came up with. Ouch.
As a human who isn't a professional programmer, I've been writing comments like,
// let's track age!!
// this is harder than you'd think as I with totally impressive
// foresight didn't add age to the raw data.
//
// More honestly, I didn't want to add age to the astro data as that's
// a calculation that can change depending on how you slice it.
//
// Hence we need to figure out their age first.
I think this passes. In general "why" over "what". Give context to why something is made like it is (when seeming convoluted or strange). Sometimes I think one can give historical facts for really hairy hard to fix issues that have seen multiple iterations. But LLMs don't see these nuances. They frequently smuggle in completely irrelevant details in comments, e.g. including details from the given task context, not understanding what is relevant for the code module as a whole.
Edit: For API comments it's "what" of course, detailing the workings and contracts of the exported method, function or type, so one doesn't have to read the code to figure out how to use it.
I'm mostly writing code for myself, but it's a project that'll end up being public and it'll be available for others to do whatever they want with. Does that change the answer?
Far too often the answer has been that it doesn't matter, because the reason you stopped using Jira is the company stopped sending paycheques.
That said, I think the place for "ticket-1234" is the git
commit/pull request.
Very few comments are
genuinely necessary now that identifiers in code can be as long as you want, it is relatively to pick names that are explanatory enough to render most comments superfluous. 1% exceptions for unusual algorithms. (You're using named consts/enums rather than magic numbers, yes?)
personally it's fine and I've thanked myself many times for overly detailed comments coming up to some from 8 years ago and thinking how tf was I so smart/stupid (depending on the context)
Claude writes comments about how things used to work, which can be useful sometimes, especially if it's a big change that requires one to genuinely consider legacy behavior, but most of the time it shouldn't be there.
Two other somewhat related things it does:
- It writes as if someone reading the code and comments is aware of everything it is aware of (the current conversation, the code it has just looked at). It's really hard to make it understand that things need to stand on their own. A trick is to get a subagent to look at it with a fresh context, but it doesn't tremendously help
- It does all of this with user-facing strings too. Claude loves to write up tooltips and other labels that leak everything to the end user. Every single concern we have, every edge case we've meticulously made our code handle, it passes on to the user, so they don't "need to worry". But no sane user would think of these things. For them, a feature is a feature. The "dynamic scheduling" button should state what dynamic scheduling does plainly, and every edge case is handled by us. The "add" button does not need a label letting the user know that they will later be able to click the "delete" button, because the user will just realize it due to our adherence to proper design. Claude fails to understand good UX for the user cannot be replaced with endless labels and explanations.
It's an uphill battle and all attempts at solving this (or the brain-dead way new Anthropic models write) usually fail to work with me.
reply