Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

A small model necessarily is missing many facts. The large model is the one that has memorized the whole internet, the small one is just trained to mimic the big one.

You simply cannot compress the whole internet under 10gb without throwing out a lot of information.

Please be careful about what you take as fact coming from the local model output. Small models are better suited to summarization.



I don’t trust anything as fact coming out of these models. I ask it for how to structure solutions, with examples. Then I read the output and research the specifics before using anything further.

I wouldn’t copy and paste from even the smartest minds, nevermind a model output.


> The large model is the one that has memorized the whole internet

This is totally wrong and a potentially dangerous way to think about LLMs. They have no clue about what's factual knowledge and what is not, per design.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: