Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
shay_ker
3 months ago
|
parent
|
context
|
favorite
| on:
Accelerating Gemma 4: faster inference with multi-...
curious that they are doing speculative decoding and not baking MTP into the model, like Nemotron
https://docs.nvidia.com/megatron-core/developer-guide/0.15.0...
zargon
3 months ago
[–]
They're using the term speculative decoding but doing MTP. It's the same thing as Nemotron, but Google removed the MTP heads from the original safetensora release. (They were not removed from the LiteRM format.)
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
https://docs.nvidia.com/megatron-core/developer-guide/0.15.0...