Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
boroboro4
73 days ago
|
parent
|
context
|
favorite
| on:
Performance per dollar is getting faster and cheap...
It's very unclear what's special in Rubin to be optimized for inference? I can see disaggregated bit (with having separate prefill and decoding nodes), but what else?
villgax
73 days ago
[–]
Lot more SMs & Tensor Cores for NVFP4 going by the looks of it.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: