Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Out of curiosity, what's currently the best model I can use locally?


With an unlimited budget, Kimi K3 (which is quite comparable to this Qwen Max imo). With a normal budget/a PC you might already have, probably Qwen 3.6 27B.


$500k - Kimi K3 (maybe $250k? Haven’t done this one)

$25k - DSv4 Flash

$4k - Qwen 3.6 35A3B Q5

$1k - Qwen 3.6 27B Q4

Some people prefer the sense over the MoE YMMV.


16 DGX Sparks can run K3 at a reasonable TPS. So that's $64k.

2 DGX Sparks can run DS4 at 1 million context with 50 TPS so that's $8k.

1 A4500 can run 35A3B. Those are about $1200 new.

27B actually takes more hardware to run than 35B because attention is done differently I believe and therefore KV Cache takes a lot of space. It will run on an A4500 but it's slow and context will be like 32k.


These numbers look about right based on my experiences as well. Though for a single user I think 2x DGX Spark (~$10k) runs DSv4 Flash fairly well right?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: