Hacker Newsnew | past | comments | ask | show | jobs | submit | WASDx's commentslogin

Likely because it uses fewer thinking tokens (that you don't see anyways).

Give a task you have to 3 different models and see what actually works for you. There are no good benchmarks.

This is like saying mass production is "cheating" against handcraft.

It is, if you are trying to sell all of them in the same market, so 99 stalls with mass produced stuff, one with handcrafted stuff, but you don't know which is which.

Once it figures out a puzzle it could probably be instructed to design a specialized harness for Luna to be able to solve other instances of the same puzzle. Minimum wage workers are not solving novel problems.

They explain it here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".


With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.

This technology is strictly an extractive parasite on the world. Use it, but don't be excited.


My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.

global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.

i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.

depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.

the work does nothing or causes net harm.



[flagged]


Would you care to discuss the topic, or just throw grenades? Surely you can come up with something more substantive than this

Ok. This requires the notion of "intrinsic value", which I believe does not exist (all value is subjective), yet is a foundation of all Marxist theory.

it's easy. people are intrinsically valuable.

do you believe that people are not intrinsically valuable? that their value is what they do for others, that it is not they themselves the person.


Buddy, admitting your thought processes forcibly terminate on pre-programmed keywords isn't a flex.

Why "terminate", Marxist philosophy is a legitimate topic, deserving to be studied. Like a rich sci-fi lore or a history of Tarot magic. Deep, fascinating, and wrong.

How dumb are you?

And yet you terminated. Which would be the correct thing to do if Marxist philosophy would be wrong on all counts, as you explicitly state. Which of course, it isn't.

You’d expect that, wouldn’t you? But, alas…

I'm using AI to build things I wouldn't (and/or couldn't) have built before.

That's the opposite of parasitic.


Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.

People already started using contributor API, and your input is irrelevant.


Don’t you have some looms to break?

I’m retired so it won’t be replacing my labor :)

The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?

I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.

3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.

i find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium

If their platform allows any kind of processing as it sounds then anyone could just dump the data. So I don't see why they would not allow downloads for local processing.

Reading between the lines, I assume they aren't reasonably capable of serving up ~TB scale data to hundreds (thousands? more?) of separate clients per day. The usual fix for this would be something like an AWS requester pays bucket.

To your question specifically, in such a scenario they would charge an obscene amount for the bandwidth so you would be strongly incentivized not to download hundreds of GB of data from them.


I think the privacy argument that keeps coming up is overrepresented. Certainly ZDR is enough for an absolute majority of use cases? I see so much talk about local inference but I doubt most of it has privacy as a valid argument (not arguing it doesn't exist). It's fun to do things locally though. I've tried it as well but cloud is just faster and cheaper.

These companies have displayed zero respect for everyone's intellectual property getting these models trained.

I think not giving them your complete trust is reasonable! I'm not saying zero trust, and ZDR is fine for most things but I understand the people who don't want to stream their whole codebase out token by token.


Then use other providers hosting open models. Companies and individuals already put their whole code base on the cloud. I'm genuinely interested in privacy-oriented use cases where ZDR is not enough.

ZDR is built on trust. Given that end-to-end encryption fundamentally doesn't work with LLMs, as they need the content to be unencrypted to operate on it[1], you have no way to prove that once your plaintext data is on somebody else's server they aren't doing whatever the hell they please with it. All you have to rely on is their pinky promise that they won't do anything with it. Trust is a valid option, much of our society runs on trust, but you can eliminate the need for trust whatsoever by running on your own hardware.

[1] Yes, I'm aware of experiments to operate on encrypted prompts, but these are only research attempts, not something that could actually be used with frontier models in production.


I'm not that worried about the codebase itself. I'm worried about the fact coding agents poke around the terminal and system so much that there is almost a certainty that some of your other personal data ends up in the context somewhere which is getting logged in to a training dataset by random hosting providers.

Privacy isn’t only, I don’t want anyone to have access to my data. It could also be, I don’t want anyone to know my use case because it’s niche and highly profitable.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: