Hacker Newsnew | past | comments | ask | show | jobs | submit | EbNar's commentslogin

Since a few months, I almost exclusively use the Chinese "flash" models for my needs. They are a joy and they cost pennies per answer. Great job.

I am legitimately more excited for this release than any frontier models at this point.

I don't need a model that can invent new mathematics. I need something that is fast, cheap, and consistent. Give me that and I can build and scale.


LLMs are not consistent

True, make them cheap and fast enough and you can scope and stack agents sufficiently that the error rate tends close enough to zero to be meaningfully useful.

Yes. I'm working in the agent industry and my god are we excited on new versions of Chinese flash models. The direct competition is Gemini Flash, and these models are much better on agentic tasks with fraction of the task price compared to Gemini. Things like oh here's a set of simple instructions for you to follow, call these tools, return this report. 20-30% of the price per task. And especially Deepseek Flash produces better quality than Gemini does.

Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message.

If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing.


Which versions of flash and at what thinking levels? Which chinese flash models and at what thinking levels? What tasks? What completion rates? How was quality evaluated?

- Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash

- Medium for Gemini, high for Deepseek.

- Things like find information, then understand something about it, then send a slack message or email etc.

- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini

- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.

Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.


Thanks... I have been trying to figure out some things. Been doing my own evals. Flash 3.8 does burn a lot more tokens on high. Interesting how smart and not smart it is. For personal use almost impossible to justify the cost of 3.8 Flash cost.

Deepseek also burns a lot of tokens, its output on high is 2x of Gemini on medium. But it's dirt-cheap so it still can be 60-70% cheaper.

From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.

All this really needs evals, the token prices tell nothing.


I am surprised at how well DSv4 flash does in the real world vs many benchmarks. You look at Flash 3.8 and it supposedly beats opus 5 and deepseek is far below.. but they were measuring efficiency, whatever that is… Something doesn’t add up for me on the published benches

[dead]


Well, it's much more than that. In general everybody's building agents now. You see these things that can help you to do things like adding things like OCR an appointment from a picture of a hand-written paper and add it to your calendar, search things from the internet, find that email with a PDF and add it to your local paperless instance.

Building an agent like this by yourself is really easy. Now, we have Gemini's subscription, OpenAI's ChatGPT subscription and all those, 20 bucks a month right?

What if you can spend that 20 bucks in tokens to do your own. And you pay 15 bucks _a year_ in tokens to run that? And you own the data, you own your code and integrations. It's really easy to do, and these flash models are _more than enough_ for simple agentic tasks.


Yeah, the latest batch of <256GB Chinese models are really nice. They're far less cryptic than Claude, and competent enough to feel almost near Opus. I canceled all my subscriptions and switched to running the Chinese models locally (not as a cost saving measure).

My mental bias always kept me away from Chinese models. Because i know that china is a surveillance state and all the things we know about CCP. But after what we learned about OpenAI and how they most likely used user data to basically cheat in an open competition i think it does not matter which AI provider you use all of them will own your data and all of them can spy on you. So I am willing to switch to Chinese models. This way we help them develop and improve models some day we can run them locally.

These models are open-weights. Anyone can host them, you don’t have to use chinese servers even though most of them offer zero data-retention policies.

>>zero data-retention policies

Yeah, that’s basically an industry-wide scam.


But it makes the compliance team happy.

The Chinese labs have released interesting papers to accompany their releases too, especially DeepSeek and Kimi. This improves their standing among a few of us, who really like to see and read the papers with details about what they have changed and how their models work.

And just a day later we have the tech report available for V4.1: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...

the US is also a surveillance state except about 80% of the surveillance is private companies (that are closely tied to the state)

Even you think both cheat, you can't possibly think they both cheat the same amount.

You're naive if you're thinking the scumbags running the US companies aren't using your data.

In any case old rules apply: if privacy is a concern don't share the data. I share all my work-related code because it's worthless, but I don't and would never share company business and process details, access to production/user data, etc.

Meanwhile I know of people connecting all the kind of MCPs for datadog/sentry/jira/concluce/production databases to their harnessess..lol.


The new meta model is fast and very cheap as well, and when used through OpenCode you get quite a lot of free tokens. But meta is also THE surveillance company, so probably also not a good choice in your case.

If you work on open source projects, I don't care about surveilance. It's right there on Github with full history anyways.

Is it? You're still putting a lot of thought and guidance into the agent's harness, the final code is just a tiny bit of that. It's like giving a junior developer final code vs explaining the whys and nuance. Which I'm not sure I want to give Meta

You can use those models from Openrouter, they have many Non-Chinese providers.

Same for me. DeepSeek models are incredibly good at implementation and light planning. I still default to Opus models for feature planning, but for most simple features the Pro models suffice.

Incredible good value and product they have built.


I recently had to config my harness to watch for cybersecurity flags from astra and funnel requests to flash when they occur because Astra gets queezy when you talk to it about UDP packets in games.

Works fantastic. Glad there is a more 'uncensored' thing to fall back to when the frontier folk are too sensitive.


The only positive side is that it is harder for students to feed university exercises to the agent in cybersecurity and expect it to make them all.

Do you use them for coding with your harness or do you use them in production? I found the latency distribution on OpenRouter to be unusable for DeepSeek v4 Flash.

So, is Mozilla promoting a search engine monoculture? Yeah, "Google bad, but..."


Or is Google the best at monetizing queries and Mozilla is using whoever is best at doing it.


I've been using it for 20 years and switched 4 yearso ago to something else. My life is fine, thanks.

EDIT: so, the downvotes are because what, exactly?


I didn't downvote you, but your post has no useful information and looks like pure Firefox bashing for no reason. You don't tell why you switched, to what, what you gained. Not even mentioning that wherever you switched, your protection against ads and tracking must have become worse, as only Firefox supports Ublock Origin.


The comment they replied to is just as useless though.


The comment to which I replied wasn't that useful either.


Does this justify adding more useless comments and asking why you're downvoted?


There's no need to make assumptions. The source code is out there. Please, show us the relevant commits which do what you said.


Ads on Brave are opt-in (and not embedded on webpages anyway) and the "crypto nonsense" is a way for them to try to stay afloat without having to accept money from Google and the likes.


They are opt-out, not opt-in. Big difference.


Opt-in. Someone else linked Brave's blog below. Read it.

https://support.brave.app/hc/en-us/articles/38305898674957-H...


That only hides them, this disables them: https://rentry.co/browserconfigs


Have you ever actually used Brave? Because then you would have noticed the huge glowing opt-out billboard of a new tab page.


It's my primary browser since a few years. If you are talking about sponsored backgrounds, good luck finding any "major" browser that doesn't have something sponsored on it (opt-out).


Let's hear it then because I'm using Brave and don't see the opt-out.


https://support.brave.app/hc/en-us/articles/38305898674957-H...

and simply dont use their homepage as your new tab page


Cheers!


Nope. Why should my money (taxes) fund a US corporation?


Why not if they are working in your interest and produce open source code.


It's not in my interest. I couldn't care less about Mozilla. It's entirely their fault what is happening to them and to FF.


Here come the downvotes. People, donate your own money to Mozilla, if you're so concerned about its fate. Let me decide what to do with my own money.


In some places it is something sadly absent or overlooked...


And with countries with lots of newcomers it can take decades (well, generations) to take on the host countries manners rather than sticking to home country manners that may not mesh with host country’s manners/customs.


Right now we have 30 ºC at nighttime + 80% relative humidity (and no, no AC). I can assure you that coffee is the last of my problems.


Also high pollen count. I am much more sensitive to heat than my partner but she has allergies. We sort of alternate. Sleeping with AC all night has its own issues.


> Sleeping with AC all night has its own issues.

What are the issues sleeping with AC all night?


Dried up air (I love A/C regardless).


Does it dry up that much? I have an ikea smart thermometer that's in the flow of the AC and I haven't seen humidity drop below 40% even in summer


To some people below 60% is too dry


This reads as a lot of fear mongering which in the end can ve resumed as "We are going to lose money".


ControlD is pretty cool.


+1 for ControlD. I've been using for a while, and whilst it does have false positives sometimes and I need to manually add a site to the allow list, it works great for me. Also, their support has been very helpful whenever I needed something.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: