I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard
Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.
I've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts.
It appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.
Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one...
Tested this model, and gpt-5.6-terra-high.
Results: this one had few issues. terra: none.
These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different.
They are good till the project is simple... not anymore.
The flash model will always use an outdated Treafik version that is not compatible with the newer docker engine, I tried to deploy some personal services with Traefik and everytime it uses this wrong version, and then fixes the version issue in the thinking chain.
I was thinking to switch to Caddy but with your experience I'm gonna stay with Traefik and bare with the version issue...
Honestly they should all use their respective pelicans as their logos. Or maybe a browser plugin to do do that on the Hugging Face and OpenRouter sites.
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
Basket? Fish? All I see is the model recursively running itself locally on an eye-pad, which for some reason beyond our understanding is obscuring the invisible fork.
This model is not very good at coding, but it is quite good at research, evaluation and action, I don't write code, but it really goes head-to-head with the most expensive models in searches such as stock market and forex
I don’t know about you but I find the information about prices, effective price (weighted average), providers and performance, benchmarks (down bottom) very useful. With openrouter I can even test it right away and compare with other models (use the chat functions).
its for the month and day the model released in 2026. 0813 -> Aug 13th. I assume that if they use it internally the padded 0 makes finding the newest model easier cause all the numbers for the date line up instead of the zig zag you get without it once you get to 2 digit months.
Welp gonna give Deepseek more money. This is very cheap indeed. And I’ve been using them and kimi for a bit now not via open router but on my own and have found them on part with sonnet 5 though sonnet 5 these days I think has gotten worse.
At work I had to move to Fable to get decent work results.
Have been letting it spin pretty hard (~$12.50 for 2B, 50% cache hits) on my traffic simulator/distributed physics engine all day, it's found some pretty significant gains without introducing any new problems.
Writing custom WMMA/MFMA kernels for an exact integer matrix multiplication library that uses RNS & CRT + Int8 GEMM to get ~90% of the theoretical i64 TOPS output from a 7900XTX (3.9 TOPS vs the 0.5 or so you get with naive hip-direct usage)
I'm not sure if you think that's a lot, but that was barely even 8 hours. I've had 200B+ months lol
79 comments
[ 0.27 ms ] story [ 15.1 ms ] threadCompetitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
I can't wait to try this new one.
Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.
Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.
Tested this model, and gpt-5.6-terra-high.
Results: this one had few issues. terra: none.
These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different.
They are good till the project is simple... not anymore.
I was thinking to switch to Caddy but with your experience I'm gonna stay with Traefik and bare with the version issue...
- https://api-docs.deepseek.com/
- https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)
Wondering, sorry if it's a dumb triviality to ask.
Is this even a (sub-)version number? I mean the major version is clearly 4.
At work I had to move to Fable to get decent work results.
I'm happy
I'm not sure if you think that's a lot, but that was barely even 8 hours. I've had 200B+ months lol