Hacker Newsnew | past | comments | ask | show | jobs | submit | uludag's commentslogin

If said game, video, or social media site used 1-2.5 kWh of electricity for every 10s of content then I'd say definitely!

Thankfully web browsing is no where near that.


You must have your units really messed up because there's no way that's right. A GPU drawing 700W for 10s would be 1.9444 Wh or 0.0019444 kWh and that's if the GPU was only serving your image gen for the whole 10s which is unlikely, most of the time is probably queuing since even a much lower powered GPU in a desktop PC can generate the same image in well under 10s.

Do you have a source for 1-2.5kWh for 10s of content? It takes about a minute or two to generate, so you'd be talking about a GPU consuming 30_000W-150_000W, it's just not possible. Even if it was distributed (which I don't think it is), that'd be 30-100 datacenter GPUs running at 100% to generate one clip? There's no way that would make financial sense.

I'm actually extremely confident that I can use the architecture to make a sweeping claim on what it can or can't do and will be extremely surprised if proven wrong:

A pure next-token language model won't be able to give detailed instructions to an ensemble of motors, mimicking a human body, to do a wide variety of tasks our human brain is excellent at doing, for example, inserting keys into a car, opening the door, sitting down, starting the car, putting the car in reverse, and exit a parking lot, being careful not to hit anything.


And what's wrong with downplaying the abilities and faculties of AI models if that's what people feel like saying? We don't call humans or animals sacks of chemicals because we believe they have moral status.

Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?

Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.


What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.

I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.

Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.


I think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI.

If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-AGI at measuring AGI and that's a bit silly.

(Edit: I wrote ARC-GIS the first time around, for some silly reason)


It's not measuring AGI at all, it starts from human "core knowledge" so it is parochial. It is made of tests that still fail so by definition next version will also start low. Moving goalpost.

This is a big reason why I feel like even though LLMs are _effectively_ AGI in some regard, they also are a hack around what most people figured AGI would look like before the advent of LLMs. Humans can do metacognition, output multimodally at the same time (verbal _and_ physical intelligence go together to produce an expressive face while one talks), have a good sense for what they do and don't know, continuously take in and respond to the world around them in a (mostly) uninterrupted fashion without "turns", learn knew knowledge and retain it for their whole lives, etc. When you reduce a human to a text generator, yes obviously SOTA LLMs perform way better, but rather than invent something that can operate as an always-running "being", we've grafted a harness around an intelligence that is bound purely to speak only when spoken to. Maybe organic intelligence is already that, playing out at a super high refresh rate, but I don't know.

Current AI is arguably much more capable of multimodal output than humans. It can produce an incredibly vast variety of audio, images, and video. Humans are limited to producing the sounds we can make with meatflaps in our throats, and contorting various parts of our bodies to produce crude symbols and shapes.

(Very capable!) Embodiment, persistent operation and continuous learning are indeed things that still set us apart from AI. None of those are fundamentally difficult to solve, though.

More importantly, none of those are particularly relevant for being "intelligent": If a criminal threatened to kill your family unless you solve some difficult problem that requires only intelligence and you could choose any single person, animal, or AI to help you with it, which would you choose? Be honest.


The thing I trust the most to solve tricky problems reliably is a specific very skilled programmer I've known for twenty years.

I wouldn't say he never makes mistakes, but his success rate is a damn sight better than any LLM I've ever interacted with (and I drive Opus daily, due to corporate demands to use LLMs).

I'd pick him.


Same partial answer as I gave your sibling commenter:

"OK, so the task ends up being to write a 40-page analysis of the result of a specific experiment in quantum chromodynamics, to be finished within 2 minutes. Does your choice for a human work, or would you have rather chosen any of the frontier AI assistants? Be very, very honest. Remember that the lives of your family are on the line and nobody will judge you for a lack of allegiance to humans."


This task is intentionally designed to ensure a human cannot do it.

The initial scenario is utterly, insanely absurd to begin with, but I tried to go along in good faith and gave you the true answer.

The result was a bad-faith rhetorical trap, so I'm done with this thread.

In another attempt at good faith, as part of bowing out I will add some actual response to your anti-useful cheap rhetorical trap:

I do not trust LLMs to get things right in high-stakes scenarios. I have seen the current models spit out falsehoods and errors regularly in the handful of fields I have expertise in, and have no reason to think they would do otherwise outside my expertise.

The scenario you describe is an absurd fiction, and no human making the absurd threat could evaluate the paper in less than hours (realistically even an expert would need days, and a nonexpert could not do it at all [short of becoming an expert]).

So, there's no point trusting a bullshit machine to save my family - it might very well get them killed, and whether it was right or not, what would actually matter would not be its correctness, but what the presumable bullshit machine evaluating my offered input spits out.

So, the best move I could realistically make would be to put a stab at prompt injection into the input.

For that job, I probably would actually prefer aforementioned programmer over any other option, come to think of it - I suspect he'd have better success than even another model (especially considering the safeguards the models no doubt have to try to keep users from using the models to inject other models).

Again - I'm disappointed in your worthless rhetorical cheap shot.

I suspect you'll have much better success convincing people LLMs are intelligent if you engage in good faith, listen to their perspective, and address their actual thoughts, instead of devising the sort of inanity that comes out of high school debate clubs, where people literally want to score points instead of find truth.


> This task is intentionally designed to ensure a human cannot do it.

It is one of a vast array of things the hypothetical kidnapper could come up with, some of which humans I agree will do better at (currently) and some of which AI will do better at. We clearly agree that in that array there is at least one task that a frontier AI would be better at than any human you could pick.

It is a thought experiment, so there is nothing fundamentally wrong with it being extreme or unrealistic (thought experiments very often are), but for the sake of goodwill let's 'weaken' it a bit: the AI or human always has an hour to come up with the answer, the question and answer are in their preferred language, and the answer fits on four pages. You can't help them, though; The criminal 'prompts' them. They can use the internet as an informational resource, but they can't communicate/ask for help/post anything (with the spirit of this being: no loophole in letting somebody else do the task or parts of it for them). And of course all subject matter of all complexity is fair game (including but not limited to quantum chromodynamics experiments).

Given that situation, do you think the programmer you mentioned would be more successful than a frontier AI in more than 50% of the possible intelligence tasks?

Edit, addendum: Please, if you can, also let said programmer read this thread and give his opinion on it. It sounds like he would have interesting things to say on this.


Before "AI," humans have created a vast array of "multimodal output" (computer art, instruments, dance, architecture, etc.). Why are you giving the AI a harness and a plethora of tools and not the human in this comparison? Without these, the LLM too would be utterly useless.

Also, that situation is extremely contrived. If a criminal threatened to kill me if I misspelled a word, I would choose a dictionary. If a criminal challenged me to predict a next token, I'd choose the LLM. For all real precarious dangerous situations, I would obviously choose a human. Like immagine the hilarity (or tragedy) that would pursuit if ChatGPT tried to handle a hostage situation or a plane hijacking.

and those meatflaps are normally called vocal folds/cords btw


> Why are you giving the AI a harness and a plethora of tools and not the human in this comparison?

I am not. Multimodal models generate that output directly, without tools. Which 'tools' does AI use to generate all those images, songs, and videos do you think?

> Also, that situation is extremely contrived. If a criminal threatened to kill me if I misspelled a word, I would choose a dictionary.

Of course it is contrived, it is a thought experiment. Does not make it less valid. It is essential that you don't know what task it is going to be, just that it is a task requiring a lot of intelligence. This way question dodging loopholes like "I'd choose a dictionary" are impossible (and people will always try to find some cheesy exit rather than facing reality). You have to commit to something or somebody that has broad and general intelligence; you do not have the luxury of choosing the perfect tool for a very narrow task.

> For all real precarious dangerous situations, I would obviously choose a human.

OK, so the task ends up being to write a 40-page analysis of the result of a specific experiment in quantum chromodynamics, to be finished within 2 minutes. Does your choice for a human work, or would you have rather chosen any of the frontier AI assistants? Be very, very honest. Remember that the lives of your family are on the line and nobody will judge you for a lack of allegiance to humans.

Again, don't go for shitty loopholes. Engage with the thought experiment in good faith and thus as it is stated, not some conveniently distorted version of it.

> and those meatflaps are normally called vocal folds/cords btw

What? Next you're going to tell me that meatspinner is also not the name for the human male reproductive organ.. Maybe I need to get a refund on my Temu Gray's Anatomy.


This is a much underappreciated point.

That said, a look at the state of self driving and the recent robot olympics shows that advancement on that has accelerated enormously, though whether it's reflected in any of the LLMs is something else entirely.


AGI has a pretty precise definition, covering only cognitive tasks.

Running a marathon is not needed to claim AGI.


>AGI has a pretty precise definition, covering only cognitive tasks.

OpenAI's own charter defines AGI as "Highly autonomous systems that outperform humans at most economically valuable work". This is actually fairly sensible and involves obviously a ton of non-cognitive, emotional, social and physical activity. In other words, if you can replace most or all human beings with a machine, you have something that's generally intelligent.

That's obviously not even remotely where we're at, AI chatbots do well on narrow usually text based or programmatic problems, but can't even replace a barista or a plumber.


There is rapid progress in generality in humanoid robotics though. I think within the next year or less we will get the ChatGPT moment for humanoid robots. If you look at progression of capabilities such as the recent Skild AI demos.

Hard disagree.

The problem is sensors.

There are simply no technologies today that can replicate the density, precision, and versatility of human touch sensors. Until then, there is simply no way to create generally capable robots that can operate at the level of a human.

And unlike LLMs, advancement is held back by physical limitations like materials science, so progress has been and will continue to be much slower.


What task do you think that humanoid robots can't do? Also, we don't need fully equivalent touch to get useful performance.

If you look you will see a really broad range of tasks accomplished already, including thing like manipulating screws, picking up pills, inserting wire harnesses, folding clothes, putting away dishes. And there are several companies with built in or component advanced touch sensors like Figure or leading edge touch sensor companies like SynTouch and GelSight.


Peel an orange? Crack an egg? Thread a needle? (Heh, drive a car...) There's a huge range of tasks that a non-specialized, general purpose robot simply cannot do yet. I'd be easier to enumerate the things they can do than the things they can't given the current state of the art.

Sure, build an orange peeling machine and it'll do great. But that's not what we're talking about here.

As for those demos videos we often see, those are very highly choreographed demonstrations. Show me a real life humanoid robot operating free form on a factor floor and doing those things and I'll be impressed.

And to be clear, this is not meant to understate what's been accomplished. I'm just saying the path for advancement is a lot harder and based in physical rather than computational limitations, which are much harder to overcome and go much slower. We simply cannot look at the growth curve of LLMs and expect robotics to advance at the same rate.


peeling an orange and cracking an egg already demonstrated. Figure 02 worked on BMW's actual Spartanburg production line for about 1,250 hours running 10-hour shifts.

keep paying attention, you will see how wrong you are about it being physical limitations as the physical AI continues to rapidly improve.


Surely part of the problem is that intelligence seems to be implicitly conditioned on embodiment, to the point that "covering only cognitive tasks" seems inherently ill-defined or arbitrary. Everything we do is a cognitive task. At some point, the criticism will be "sure, it can solve research-grade math problems, but it can't fold my laundry".

Even our large language models have an implicit embodiment in the domain of text (and more recently, multimodal inputs). That seems sufficient for certain things, and insufficient for others. I suspect that AGI that does everything a human can do eventually turns out to be fairly analogous to humans in terms of sensory input and domain output, even if the scale is radically different (e.g. thousands of robots uploading (touch, sight, audio, smell, etc.) sensory data to a single model, and each being actuated individually).


> can match or exceed human cognitive abilities across a wide range of tasks

If you go by definition AGI is not general, just "smart ape" shaped.


I can just immagine the response to a headline "LLM chooses mass death: thousands killed in horrific AI accident" being something like "lots of humans have caused mass death too."


> So I vote we just ban almonds. Problem solved!

I'm pretty sure if we had an actually vote between banning AI data centers and banning almond farming, one would win by a landslide. I'm actually confident that if some hypothetical community had absolutely no water or energy problems they would vote against having the AI data center built.


I think the point is without taste you wouldn't even know what things should be copied out of the countless competitors. How would you even know in the first place if a competitor did a particular thing tastefully without a good sense of taste yourself?

Also, there are a lot of areas where taste will absolutely reign supreme (and has been), especially anything related to media or entertainment (books, articles, videos, games, music, etc.)


Isn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens.

Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many tokens as it can doing the wrong task. This happens to me all the time.

Or if I was to tell a model to do something and it realized it could only do it by hacking into another service, by all means it should ask me if I want it to hack and not go ahead and do it by itself.


It's likely bumping up against people's desire to have the model complete a given task without asking the person to intervene a bunch of times.

Seems unclear how you satisfy everyone here.


Provide an option. An autonomous mode or an interactive mode.


Give the model judgement


Whose judgement?

I'm not a fan of how many "I set budget X and woke up to an eleventy trillion dollar bill" posts I see, and those are all generated by companies giving their tools the 'judgement' that the completion of the task is more important than the users wallet (or more cynically that they think they can actually get paid by just blasting unintended compute)


And good taste.


> And good taste.

There goes the ability to use the web as a training set.


And common sense.


If only two people could agree on what that was.

Seriously, the only time I see people talking about "common sense" is to criticise its perceived absence. Every example people have given of "common sense", not just to me in my life but historical examples reproduced and passed down over the millennia, has been false in some important way where treating it as true held humans back for ages.


And then finally software engineering work can be completely automated :).


You misspelled “Invent AGI”


+1 Opus 5 is more creative and broad which is great. But I feel the new "use your judgement" leads [1] to it often escaping its sandbox or cheating, it keeps breaking specific rules I explicitly set. That eagerness [2] may be desired for agentic coding but it really breaks writing specifications, documentation or research and yes burns tokens doing unasked for work.

[1] https://claude.com/blog/the-new-rules-of-context-engineering... [2] https://simonwillison.net/2026/jun/11/fable-is-relentlessly-...


One of the key limitations of the last generation was that they biased towards inaction and gave up. (Hence Ralph-loops.)

Seems pretty clear that most people wanted these models tuned to “bias towards action”.

It’s on you to set the /goal and prompt context such that it asks you for input on what you want to be consulted on, and only brute-forces the parts of the problem that you want it to.


Is there any recommended standard copy/pastable snippet to tell the model to ask, if the user’s input could mean an easier approach?


I have been using https://www.aihero.dev/skills-grill-me with great success. It’s a very nice small skill for lightweight up-front planning. And if you have that at the top of your session, you can just say “please use AskUserQuestion for any important decisions that arise which might alter our plan”. It seems to prime the agent to be in high-level discourse mode.


I don't think it is. It's really easy to get a persistent, clever, hacky model to dial that down a bit and just come back to chat before stomping off into the woods.

If a model couldn't ever do that in the first place, it'll just get stuck.

I work in the "ZeroOne" space, working on concepts and prototypes for things that don't exist in market yet. Sometimes these models crank hard and immolate tokens while grounding themselves on expensive-to-ingest self-developed frameworks. If the results are well judged and the crank-turn latency is low, I'm okay with the cost as long as the model isn't wasting my time.

But when I want to do more boilerplate work, I turn down the model and thinking level and get more traditional about restraining action. For the really hard stuff, I reach for the models that will start a token bonfire in the back yard.


I don't see why? If they give you a math test and tell you you cannot use a calculator, should you just say "please can I use a calculator" and quit?


Cut down that tree.

Can I have a saw?

No

Okay, it will take much longer then as I’ll have to do x, y, and z.

Pick one: [That’s fine, proceed] [Okay you can use a saw]


Preferably, it would ask for help to see the image and only engineer its own workaround if that was not an option. There could also very well be something in the prompt that would cause it not to do so. Opus 4.8 is usually good at asking for clarification for things like this when I'm using it.


Why would it ask to see the image if it was already told it cannot see the image?


No, a proper analogy is painting with blindfolds on or playing the piano deaf. Doing math without a calculator is like… rather common?


That violates the golden rule of ai economics: always choose the path that burns more tokens


another way to describe this is we as the user should get to limit tokens per prompt in a way that prevents these kinds of runaway solutions. It’s like if you asked your Lead Engineer to reduce build times and didn’t give him any budget restrictions. The next month he reports he reduced build times by 95% and you also have a $10,000,000 AWS bill.


Sometimes its worth tracking usage optimization and raw abilities as separate qualities.


Very weird indeed. People must not realize that you can completely change the response you get back from an LLM by how you ask questions. Any bias can implicitly be implanted in the question you ask and drastically modify the response. This is what I got Gemini to say about the article:

  The author’s tone in this piece can be described as brutally candid, deeply relieved, and unapologetically sarcastic.
very different from "The overall tone is deeply personal, cathartic, biting, and polemical, with flashes of humor and a deliberate attempt to soften the ending."


Not very different, mind you.

I don't really agree with using LLMs to do this but it correctly identified the attempt to soften the ending, which is to my mind significant in the whole piece; this person wants to repeat and frame unkind things he's heard, say unkind things, and then assert that he wasn't doing either.


Personally I don't think that matters, because the article is problematic enough when it can be read like ad hominem. Assuming that the question was phrased reasonably neutrally (but not necessarily free of hidden biases), the fact that LLM concluded so is an enough evidence here. Also as a non-LLM data point, I felt roughly same (especially the "softening" bit).


this is what worries me. I have friends that love what AI tells them about their personal pet 'thing' and how awesome it is. Yet not one of them has even tried once to get the AI to criticise it's own answers, and hence learn that you can trivially get an AI to make a convincing-_sounding case for any point of view.

I tell them to try, and they laugh at me as they roll their eyes and waffle on about 'tricking' the AI like its some kind of hacking.


Well paid expats wearing a hijab, who definitely aren't refugees, will not be treated nicely in Germany. Lived and worked in Germany and saw it a lot. It's a low bar indeed to treat skilled labor coming to your country nicely that sadly Germany can't even pass.


You are kind of expected to adjust yourself to the culture of the country you’re immigrating to. If you fail (or refuse) to do so in the most obvious/visible ways possible (like clothing), I think you shouldn’t be surprised when people look at you weird because you look out of place.


Someone wearing a hijab looks out of place in the land of the Döner?


> Someone wearing a hijab looks out of place in the land of the Döner?

The history with the Turkish gastarbeiters is a complicated one. Please don't twist the knife in the wound.


How is this twisting the knife in the wound? I wanted to say that they (and other people wearing hijabs) are part of German culture.


This is very bad faith and you know it.


It was in fact not, hence why I asked :/


I don't think even the women wearing the hijabs would consider it to be part of German culture.


I don't know about you, but I don't treat people worse just because they dress differently.

And similarly, I don't think it's reasonable to suggest OP is the wrong one for being treated worse for wearing clothes from their culture.

If you look out of place, you will get more looks, and that's it, that's where the line should end.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: