For maths, it's also wrong most of the time. Generated stuff looks right, then let Opus audit and see the disaster.
DS4 is usable iff you have a way to test the generated stuff, and to convince yourself that its production is right. With successive review-fix rounds, it's obviously way more reliable too, but that can't compensate for it's lack of rigor when reasoning.
It's very smart but neither rigorous nor careful. And that is a direct result of its architecture.
I can get by working on code strictly in GLM. I can't with DeepSeek. It makes some pretty careless mistakes and isn't a very deep thinker.
It is very useful as a general purpose model for non-coding purposes though.