Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

They can just train another one.


Llama 2 end weights are public. The data used to train it, or even the process used to train it, are not. Google can't just train another Llama 2 from scratch.

They could train something similar, but it'd be super weird if they called it Llama 2. They could call it something like "Gemini", or if it's open weights, "Gemma".


The article says they used maxtext to load the weights and pretrain on additional data. It looks like the instructions for doing that are here: https://github.com/AI-Hypercomputer/maxtext/blob/main/gettin...


They don't mean literally LLaMA. They mean a model with the same architecture.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: