Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It helps to be able to run the model locally, and currently this is slow or expensive. The challenges of running a local model beyond say 32B are real.


Ye the compressed version is not nearly as good.

I would be fine though with like 10 times the wait time. But I guess consumer hardware need some serius 'ram pipeline' upgrade for big models to be run at crawl speeds.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: