It's hard to believe that a model can nowadays solve mathematical challenges and break ciphers, yet it fails to do trivial tasks involving critical thinking, having taste, and not just running around in circles.
A dead giveaway that I think it's AI-made (not necessarily a con) is the complexity of the UI and redundant information. Things like the green dot paired with "interface is up". It looks polished though.
I'm sorry this makes it seem like I didn't do my research. I did a TON. To fix it I'll add a benchmark/comparison table. Also, I wouldn't call it market research since this is not commercial AT ALL.
Isn't this cheating? Or rather, are frontier agents only looking at one question at a time? If I understand correctly, you're looking at all the examples of the exam questions. If the exam was adjusted so that you can only look at one question at a time, you won't get 44% anymore.
Why would that be cheating? That's what humans do when they learn, they look for the signals and patterns that reduce the possible set of answers so they can converge on the solution and narrow the search space.
I mean, I'm interested to know if the frontier models also get to see all questions at once. Then it's more fair game than if they just see one question at a time.
reply