Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> This is important because local llm rarely has parallel streams to batch together.

I think most people using agent-like usage could easily run any number of parallel streams pretty often, but you run out of vram for multiple KV caches, unfortunately.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: