79 comments

[ 0.27 ms ] story [ 15.1 ms ] thread
(comment deleted)
I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard
Benchmarks:

    | Benchmark                | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2   | Kimi-K3   | Opus-4.8  | Fable 5       |
    |                          | 0813      | 0731        | Preview   | Preview     |           |           |           | (w/ fallback) |
    |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
    | HLE (wo/w tools)         | 42.7/60.0 | 37.8/51.5   | 37.7/48.2 | 34.8/45.1   | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0     |
    | Terminal Bench 2.1       | 87.9      | 82.7        | 72.1      | 61.8        | 81.0      | 88.3      | 85.0      | 88.0          |
    | NL2Repo                  | 61.5      | 54.2        | 38.5      | 39.4        | 48.9      | -         | 69.7      | -             |
    | Cybergym                 | 83.3      | 76.7        | 52.7      | 38.7        | -         | 80.0      | 78.3      | 83.1          |
    | DeepSWE                  | 62.7      | 54.4        | 12.8      | 7.3         | 46.2      | 67.5      | 58.0      | 70.0          |
    | Toolathlon-Verified      | 74.1      | 70.3        | 55.9      | 49.7        | 59.9      | 76.5      | 76.2      | 77.9          |
    | Agents' Last Exam        | 25.7      | 25.2        | 16.5      | 15.8        | 23.8      | 27.6      | 25.7      | -             |
    | AutomationBench (Public) | 31.8      | 25.1        | 12.8      | 10.8        | 12.9      | 30.8      | 27.2      | 29.1          |
    | DSBench-FullStack        | 71.1      | 68.7        | 41.8      | 37.0        | 61.8      | 73.7      | 71.6      | 77.2          |
    | DSBench-Hard             | 67.2      | 59.6        | 31.1      | 25.8        | 54.5      | 63.0      | 71.7      | 68.3          |
Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...
Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
Just use opencode go, you get more bang for your buck. Same api
What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.
Opus 5 fucking sucks to talk to and read compared to 5.6 Sol though. I’m fully done with Claude models until they figure this out
I've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts.

I can't wait to try this new one.

Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)
Still behind Kimi-K3 in almost half of the benchmarks
Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project.

Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.

Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

It appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.
Even though cost-per-token is low, Deepseek v4 tends to burn an immense number of tokens to accomplish tasks.
I can go for days on end without topping up my deepseek account. When it can’t solve a problem I switch to GPT and have to top up in real time.
Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one...

Tested this model, and gpt-5.6-terra-high.

Results: this one had few issues. terra: none.

These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different.

They are good till the project is simple... not anymore.

The flash model will always use an outdated Treafik version that is not compatible with the newer docker engine, I tried to deploy some personal services with Traefik and everytime it uses this wrong version, and then fixes the version issue in the thinking chain.

I was thinking to switch to Caddy but with your experience I'm gonna stay with Traefik and bare with the version issue...

I agree. The “frontier level” open weights models benchmark really well but fall behind in real world performance
Before DeepSeek-V4-Pro-0813's price goes up, I expect a surge of frantic traffic — hope the servers can hold up.
DeepSeek raised their prices. Oh my god.
Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Your tool is giving "Error: Gist API returned 403"
Honestly they should all use their respective pelicans as their logos. Or maybe a browser plugin to do do that on the Hugging Face and OpenRouter sites.
I assume this isn't watermarked...
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
I like to do something outlandish like a "cockatiel driving a UFO on it's way to austrailia"
Not bad but the left foot is still in the wrong place
we hitting singularity levels of bicycle chain here
(comment deleted)
Wonder if we'll ever see optimizations for pelican riding a bicycle svg, make it in to model training runs.
Basket? Fish? All I see is the model recursively running itself locally on an eye-pad, which for some reason beyond our understanding is obscuring the invisible fork.
Might be time to also have it try more three dimension pelican rendering. Or a short animated version (even just a few frames) still in standard SVG
This model is not very good at coding, but it is quite good at research, evaluation and action, I don't write code, but it really goes head-to-head with the most expensive models in searches such as stock market and forex
I'm Satisfied with this model (in opencode)
Why does this link to OpenRouter, which has no useful information on its own? Linking to the official API or the benchmarks would make more sense:

- https://api-docs.deepseek.com/

- https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)

Don’t you find the official document website was out of service for a long time since the new model was published soon?
I don’t know about you but I find the information about prices, effective price (weighted average), providers and performance, benchmarks (down bottom) very useful. With openrouter I can even test it right away and compare with other models (use the chat functions).
Again, I will wait until there's a provider that doesn't train on prompts before I will benchmark.
Is having padded version numbers with a leading zero a common thing?

Wondering, sorry if it's a dumb triviality to ask.

Is this even a (sub-)version number? I mean the major version is clearly 4.

its for the month and day the model released in 2026. 0813 -> Aug 13th. I assume that if they use it internally the padded 0 makes finding the newest model easier cause all the numbers for the date line up instead of the zig zag you get without it once you get to 2 digit months.
These people can't version control properly, V4.1 or V5 would be more appropriate.
Welp gonna give Deepseek more money. This is very cheap indeed. And I’ve been using them and kimi for a bit now not via open router but on my own and have found them on part with sonnet 5 though sonnet 5 these days I think has gotten worse.

At work I had to move to Fable to get decent work results.

Have been letting it spin pretty hard (~$12.50 for 2B, 50% cache hits) on my traffic simulator/distributed physics engine all day, it's found some pretty significant gains without introducing any new problems.

I'm happy

Can you explain how you used 12 billion tokens to do useful work?
Writing custom WMMA/MFMA kernels for an exact integer matrix multiplication library that uses RNS & CRT + Int8 GEMM to get ~90% of the theoretical i64 TOPS output from a 7900XTX (3.9 TOPS vs the 0.5 or so you get with naive hip-direct usage)

I'm not sure if you think that's a lot, but that was barely even 8 hours. I've had 200B+ months lol