Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You don't actually need 480GB of RAM, but if you want at least 3 tokens / s, it's a must.

If you have 500GB of SSD, llama.cpp does disk offloading -> it'll be slow though less than 1 token / s



> but if you want at least 3 tokens / s

3 t/s isn't going to be a lot of fun to use.


beg to differ, I'm living fine with 1.5tk/sec


Spec decoding on a small draft model could help increase it by say 30 to 50%!


i'm not willing to trade any more quality for performance. no draft, no cache for kv either. i'll take the performance cost, it just makes me think carefully about my prompt. i rarely every need more than one prompt to get my answers. :D


Speculative decoding doesn't change output tokens.


Draft model doesn’t degrade quality!


I beg to differ, especially when it comes to code.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: