100 comments

[ 3.5 ms ] story [ 22.5 ms ] thread
DeepSeek V4 Flash (Preview → 2026-07-31)

• Terminal Bench: 56.9 → 82.7 (+25.8)

• Toolathlon: 51.8 → 70.3 (+18.5)

Compared to GPT-5.6 Terra:

• Terminal Bench: Flash 82.7 vs Terra 78.4

• Toolathlon: Flash 70.3 vs Terra 53.1

• DeepSWE: Flash 54.4 vs Terra 69.6

• Agents' Last Exam: Flash 25.2 vs Terra 50.4

Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

hope that deepseek become better
Starting to wonder if the free big pickle model on opencode has been DSV4F0731 for the past few months. It’s been incredibly fast and good.

  > Whatever capabilities they get, can be used "forever" going forward.
"Forever" gets the scare quotes because it is implied only up until the Butlerian Jihad?
distilling K3 into DS4 flash will likely only be a good idea for specialists. the difference in model capacity is otherwise too large.

we used to distill GLM 5.2 into Qwen 27B specialists with great success.

developing the workflows is tricky though. we had the advantage of a straightforward mapping function in mind: English -> SQL which was relatively easy to pipeline training for.

In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.
(comment deleted)
They're always very understated in their update descriptions. This is actually a HUGE improvement in the model's capabilities rather than just a small tweak.
I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng. You have to remember DeepSeek v4 flash even though a bit cheaper, does not have vision abilities, which is a big draw for agentic tasks.

I admire DeepSeek's openness, but even they have been raising prices after their discounts.

According to the leaked call transcript, DeepSeek is working on vision for V4. Not sure when it will land though.
Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.
Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.

If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

Not just 200B model, it is only 160GiB.
MiniMax's another lab that's known for relatively smaller models (their latest, M3 is 295b) that punch way above its weight.
Wonder how good the proper version of V4 Pro will be.

I'm still considering pulling the trigger on the annual subscription of Kimi for K3 but it's sometimes slower than I'd like (at least when compared to Anthropic) even on their Vivace plan, and the token limits on the GLM Coding subscription for GLM 5.2 were too easy to hit.

Oh my goodness what an update. I need these weights. It's an incredible model for the size. The improved tool calling etc. should be able to make my harness way simpler. This runs at mega-speed on prosumer hardware (2x RTX Pro 6000).
How many tok/s are we talking about?
The old one was 160 tok/s. With DSpark it’s like 250 tok/s or higher. I don’t recall from tbe last time I was looking. You can turn on farther drafting and get more.
[dead]
If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the price decrease.

Crazy.

On DeepSWE Deepseek is 54.4% and Luna is 67%
OpenAI must've known this was coming, hence the Luna price drop. This competition is amazing!
Every time I want to have fun coding something with natural language processing, I use deepseek flash. It's just incredible for the price. I have a fairly popular app with 400 users that uses DeepSeek in the background and it still didn't hit even 50 bucks of usage in a month.
Mind sharing the name? Sounds interesting.
Finally have a model with usable intelligence, at a reasonable price. Can't imagine what Pro GA would look like, considering pro preview has only 1.6t parameters.
For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning falsified logic or information cooked in by the developers, rather than the training data speaking for itself
The previous V4 version wasn't called “Preview” by most inference providers. For example, the OpenRouter model slug was `deepseek/deepseek-v4-flash`. So now there will be confusion when someone talks about V4 Flash or when someone offers V4 Flash inference.

Why not call it V4.1?

DeepSeek being DeepSeek. v3 and R1 went over the same and had multiple versions
Very promising. So it will both keep the speed and reduced price, yet exceed performance of the quite sufficient deepseek-v4-pro?

Should be extending the lead in intelligence/cost index, as deepseek-v4-flash already were the most price efficient model, which now becomes even better. Although, in the deepseek APIs, the cost is leaking all information about codebases to China.

Supposedly better than GLM 5.2 according to at least one benchmark.
The flash variant is on par with Sonnet 5 on DeepSWE (54%). Big, if true.
How's their performance in English prose? We are currently searching for cost effective ways to keep story wikis up to date.
I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast.

I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The rest 10% is to spot bugs, security problems and to investigate better architecture, which flash can also do pretty well, I just cross check it.

Faster iterations are way better for me, I hate waiting for 5-10 minutes on small changes. I tried to use recent versions of Kimi and GLM, but they use too much thinking for no reason and are pretty slow because of it. I also often feed a lot of data to it, without worrying about hitting the limits: dependencies (to find bottlenecks in them), logs, performance dumps and so on.

Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

This is my experience also: DeepSeek v4 flash is good enough for most of my work and I like the fast response times. I buy tokens mostly from FireWorks.ai in the US, but I also prepaid for a large chunk of tokens directoy with DeepSeek.

I use OpenCode mostly (uses fewer tokens than Claude Code) and I am looking forward to the release of DeepSeek’s own coding harness.

Same experience here. Plugged it into my harness instead of a 10x more expensive model and it... just worked, sometimes even better than the expensive model in a direct comparison. Starting to think there is more value in speed than "intelligence" at this point.
Kimi K3 (instead of Opus) for expensive stuff, DSV4 Flash for tasks (instead of Sonnet)?

Does this make sense?

Judging by the openrouter leaderboard ranking for today, it looks like Dv4F us more popular than mimov2.5.

https://openrouter.ai/rankings?view=day#leaderboard-table

These days cost per task is more important, and SOTA models have become expensive.

MiMo, the overlooked sidekick to the hero. Will be interesting to see what Xiaomi bring to the table.

These massive jumps in cheap models, is really great times!

What's the best way to run this on a 64GB M2 Pro?
DS4Flash has 284B weights. 64GB? No go.
I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days:

- Cost: $4.55USD

- API requests: 3,467

- Tokens: 323,183,886

And as an engineer who leads a small team, I have very high standards for quality, and these carry across to my personal projects where I use deepseek. It has not disappointed at all for coding or review tasks. For everything else, use another model.

But doesn't it hallucinate a lot? Does that affect your workflow?
What harness are you using to achieve that level of token caching?
Can you give more info on how you use/prompt those LLMs for code review and what kind of prompts you use?

I've had worse experiences doing it because the quality of answer has been quite bad, and I'm wondering if my methods are the reason.

You are using DeepSeek's services directly? Doesn't that end up sending at least snippets/chunks of code to a server where it is subject to Chinese government data access laws? Even if I was okay with that, my organization would never be. And even if they were, our partners/vendors/customers would not be. I think that's the sticking point for a lot of people.
DeepSeek is amazing, they are, from a cost/benefit literally an order of magnitude or more better than the 'SOTA' models, and yet no one really talks about them.

I'm using them for my micro-saas, and they have made my niche economically profitable where as SOTA models are only slightly better for massively increased expense. Its truly impressive.

Word of advice to anyone, not all your use of LLM tech needs to be code/dev work related.

We are entering 'Web 4.0 era' or whatever you want to call it. Massive transformations of nearly every single business will and are being developed as the cost of intelligence as a commodity is falling through the floor...

Those are rookie numbers. Last 13 days including today, so effectively 12 days of usage:

- $19.27 USD - API requests: 7,877 - Tokens: 2,116,598,952

Ok this is a bit of lie, a lot of my tasks are very experimental loops whose 99% output is like rubbish and can work forever continuously and take advantage of that 120x cheaper input cache. Still incredible.

What are the tasks where it falls short?