Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

They have had the best math models for about a year most folks just didn't know about it. You can't find inference on APIs, but I run these at home, this is also the advantage of open models.

https://huggingface.co/deepseek-ai/DeepSeek-Math-V2 https://huggingface.co/deepseek-ai/DeepSeek-Prover-V2-671B



You are of course specifically referring to the math optimised models, not the chat ones folks would generally encounter. Not that I’m trying to contradict you, your point is super valid and I agree with you! But I’m supplementing to help anyone following along who may make choices.

This is when it happened for anyone interested: https://binaryverseai.com/deepseek-math-v2-benchmarks-review...


Shouldn't one use e.g a Wolfram Alpha MCP endpoint for math in AI? From what I've seen on even premium non-quantized models, I would never ever trust the innate ability of a LLM to calculate.


You run a 671B model at home?


Yes, and plenty of others do too. Quantizied. Join us at r/localllama

My largest models

   318G    /llmzoo/models/Qwen3.5-397B
   377G    DeepSeekv3.2-nolight
   380G    /llmzoo/models/DeepSeek-V3.2-UD
   400G    /llmzoo/models/Qwen3.5-397B-Q8
   443G    DeepSeek-Math-v2
   443G    DeepSeek-V3-0324-Q5
   522G    /llmzoo/models/GLM5.1
   545G    /llmzoo/models/kimi2.6
   546G    /llmzoo/models/KimiK2.5


Is your house's heating system based on H100s?


What hardware do you use?


I think the answer to this is:"yes"


Most of those have custom quants for Mac Studio M3 Ultra 512GB. You'll typically see them mention it by name.

All of that list but the last three run at these sizes. For last three, look for a custom quant, e.g. 9.5 bits and/or the Ultra M3 512GB mention.

Not sure which direction I'm surprised but Macbook Pro M5 Max ticks over models at the same speed. With "only" 128GB look for models of 116 GB (the absolute max that retains reasonable stability) or less.


a Beowulf cluster of 256 x Raspberry Pi 3.


I used to maintain a 2000 pi 4 cluster, before LLMs were relevant, with around 6gb free ram per node. I wonder what I could have done with something like this.


All of it.


even quantised, those are HUGE


It's a big house.


Maybe if there was a 1-bit quant.


Apple briefly was selling Mac studio with 512 GB of unified ram, meaning all that was available as vram.


There is an IQ1_S quant, but even that one is ~184GB (download size).


Vertex AI has had deep seek available via API for a while


I'm talking about their specialized math models, not the general model.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: