I don't understand how DeepSeek can be so cheap with their cache pricing - ~0.003 usd / 1Mtok. 100x less than Kimi K3, or similar numbers against pretty much any other decently sized model to my knowledge. I've been using it whenever possible as even longer agent sessions cost few cents.
If you read DeepSeek's papers, you'll find a litany of architectural features that allow for a greatly reduced cache hit price by shrinking the size of the KV-cache.
Many of these techniques haven't been published very long ago - it often takes a good 6-8 months for techniques to percolate. But also, they come at a complexity cost and, seemingly, also at a stability cost.
Also potentially a performance (in terms of output quality) cost. DeepSeek is cheap on a per token basis but lags behind in the benchmarks, perhaps it was a calculated tradeoff.
I don't understand how DeepSeek can be so cheap with their cache pricing - ~0.003 usd / 1Mtok. 100x less than Kimi K3, or similar numbers against pretty much any other decently sized model to my knowledge. I've been using it whenever possible as even longer agent sessions cost few cents.