I finished benchmarking[0] it, but it was not fun, it only supports (max) reasoning and the model is quite slow. Apart from a few requests timing out, it also has some issues with tool calling/response format schemas (Moonshot rejected tools.function.parameters with anyOf schema).
It also, for some reason failed to generate either of the 2 coding demos (hamster svg and solar system css animation).
Intelligence-wise, it's between GPT-5.6 Terra and GPT-5.6 Sol. It's ~30% better than Kimi K2.6, but a lot slower and more expensive.
It also, for some reason failed to generate either of the 2 coding demos (hamster svg and solar system css animation).
Intelligence-wise, it's between GPT-5.6 Terra and GPT-5.6 Sol. It's ~30% better than Kimi K2.6, but a lot slower and more expensive.
[0]: https://aibenchy.com/compare/moonshotai-kimi-k3-max/moonshot...