Hacker Newsnew | past | comments | ask | show | jobs | submit | docheinestages's commentslogin

It's hard to believe that a model can nowadays solve mathematical challenges and break ciphers, yet it fails to do trivial tasks involving critical thinking, having taste, and not just running around in circles.

Pattern matching is like that. It can map a full sentence to its meaning, without ability to map one of the words in the sentence to its meaning...

I disagreee big time.

You’re not really understanding how the tech works if you find it hard to comprehend.


i think the "having taste" part is more an issue with the people who use AI and what they use it for, than AI itself.

Nice try, Dario.

You must be new here...

Same here.

A dead giveaway that I think it's AI-made (not necessarily a con) is the complexity of the UI and redundant information. Things like the green dot paired with "interface is up". It looks polished though.

I confirm that I vibe coded it.

Now I'm starting to doubt the credibility of Artificial Analysis.

My gut feeling tells me it has something to do with Cloudflare. Along with AWS, they're two of the main suspects in such incidents.

…and it’ll involve BGP routing.

I thought OpenAI famously used Azure due to their partnership with Microsoft?

They use a lot of compute providers now, but they use Cloudflare for their networking

All I see is slop.

Is this a compressed and retrained version of GLM-5.2?

It's what happens when you don't do market research.

I'm sorry this makes it seem like I didn't do my research. I did a TON. To fix it I'll add a benchmark/comparison table. Also, I wouldn't call it market research since this is not commercial AT ALL.

Hey Carlos, thanks for sharing with the community! Appreciated

thanks!

Does a painter check to make sure that a portrait hasn't been painted? What a dismissive comment.

Half the people on here are using Ollama. No one is doing market research.

Isn't this cheating? Or rather, are frontier agents only looking at one question at a time? If I understand correctly, you're looking at all the examples of the exam questions. If the exam was adjusted so that you can only look at one question at a time, you won't get 44% anymore.

Why would that be cheating? That's what humans do when they learn, they look for the signals and patterns that reduce the possible set of answers so they can converge on the solution and narrow the search space.

I mean, I'm interested to know if the frontier models also get to see all questions at once. Then it's more fair game than if they just see one question at a time.

They look at each problem individually.

As the ARC AGI guidelines [1] state, 'a core design principle of ARC-AGI is that the test taker must not know what the test will be.'

[1] https://arcprize.org/policy#dataset-security


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: