DeepSeek-V4-Flash Update

636 points - today at 6:08 AM

Source

Comments

NitpickLawyer today at 6:25 AM
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

f311a today at 7:46 AM
I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast.

I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The rest 10% is to spot bugs, security problems and to investigate better architecture, which flash can also do pretty well, I just cross check it.

Faster iterations are way better for me, I hate waiting for 5-10 minutes on small changes. I tried to use recent versions of Kimi and GLM, but they use too much thinking for no reason and are pretty slow because of it. I also often feed a lot of data to it, without worrying about hitting the limits: dependencies (to find bottlenecks in them), logs, performance dumps and so on.

Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

lionkor today at 8:08 AM
I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days:

- Cost: $4.55USD

- API requests: 3,467

- Tokens: 323,183,886

And as an engineer who leads a small team, I have very high standards for quality, and these carry across to my personal projects where I use deepseek. It has not disappointed at all for coding or review tasks. For everything else, use another model.

kmarc today at 8:16 AM
Essentially I'm running everything on flash now inside pi. With the correct set of MCP servers, context reducer tooling and skills it can implement any task I throw at it. Some sessions take 30+ turns, but it's fast and cheap; all this in an hour, with ~$0.5 cost.

(TBH though, in my multi-subagent workflow I do use other, more expensive models for planning, reviewing, oracle-ing)

I haven't used our slow opus subscription for weeks.

(Also set up an OpenWebUi self-hosted chat that works from my phone, has some mcp and skills. fully replaced perplexity. Monthly cost ~$18 for hosting and subscriptions)

wkcheng today at 7:30 AM
If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the price decrease.

Crazy.

dannyw today at 9:12 AM
Deepseek and moonshot are the only two providers I consent to training for.
ggcr today at 7:07 AM
Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.

If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

heyalexej today at 7:25 PM
Long story, I have humongous zai GLM 5.2 token budget that I'm using in a similar fashion as many comments explain here. GPT 5.6 or Fable 5 for planning, GLM for implementing, researching, extracting and many other tasks I consider grunt work. Very happy with the performance, speed isn't all that good though. I'd be curious to hear from someone who works with both, DeepSeek and GLM side by side.
Goranek today at 7:47 AM
Kimi K3 (instead of Opus) for expensive stuff, DSV4 Flash for tasks (instead of Sonnet)?

Does this make sense?

thirtygeo today at 7:33 AM
For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning falsified logic or information cooked in by the developers, rather than the training data speaking for itself
baalimago today at 7:37 AM
Very promising. So it will both keep the speed and reduced price, yet exceed performance of the quite sufficient deepseek-v4-pro?

Should be extending the lead in intelligence/cost index, as deepseek-v4-flash already were the most price efficient model, which now becomes even better. Although, in the deepseek APIs, the cost is leaking all information about codebases to China.

wg0 today at 6:01 PM
Don't know about the bench marks but I am getting Opus 4.7 level performance at fraction of cost with DeepSeek V4 Flash set to high. It is a reliable workhorse.
sim04ful today at 8:42 AM
Why didn't they increment the version as atleast a patch update
namuol today at 5:51 PM
Can someone please explain how these models aren’t just fine tuned for benchmarks? I’m not plugged in to this space much but it seems like such an obvious problem…
yewenjie today at 10:29 AM
What was preventing them from calling it v4.1-Flash to distinguish it better?
throwa356262 today at 1:01 PM
troglodytetrain today at 6:47 PM
This is very exciting, my own niche micro-saas has already been able to make heavy use of DeepSeek-V4-Flash for my use case, I am looking forward to seeing how performance improves.
alecsm today at 12:43 PM
I've been using DeepSeek Pro for a while and Flash only for certain dumb tasks where I only need the speed of a LLM and not big brains.

I find the newest OpenAI and Anthropic models to be way better for big tasks that require many decisions but I don't like that anyway because I lose track of what's being done.

Knowing what I want for every prompt makes DeepSeek Pro the best LLM for me. It allows me to work relatively fast at a very low price.

sqemo today at 9:52 AM
DeepSeek is great for tasks and software I already know well. Even if it gets something wrong, I can usually verify it myself. But when I'm working with a programming language I'm not familiar with, I prefer using Codex or Claude.
arjie today at 7:19 AM
Oh my goodness what an update. I need these weights. It's an incredible model for the size. The improved tool calling etc. should be able to make my harness way simpler. This runs at mega-speed on prosumer hardware (2x RTX Pro 6000).
mordae today at 8:50 AM
I was just using it when it landed. It started reasoning more extensively from nowhere and precision went up a lot. It also changed its prose style for the better. Looking forward to weights.
nickandbro today at 6:59 AM
I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng. You have to remember DeepSeek v4 flash even though a bit cheaper, does not have vision abilities, which is a big draw for agentic tasks.

I admire DeepSeek's openness, but even they have been raising prices after their discounts.

amelius today at 10:32 AM
Note: if you are having success with a model, then please post what you are using it for. Writing HTML/CSS is very different from writing Rust/C++ or doing maths.
Reubend today at 6:57 AM
They're always very understated in their update descriptions. This is actually a HUGE improvement in the model's capabilities rather than just a small tweak.
f6v today at 8:28 AM
I'm thinking of using ChatGPT for making plans and V4-Flash for execution. Does anyone have good advice on pairing different models?
ilaksh today at 11:14 AM
I wonder when the antirez/ds4 group will have an update to their high accuracy 2 bit quant.

Although it's funny that I am thinking about that at all because I have a 2060 :P . My local inference is playing with Gemma 4 E2B and MiniCPM 5 1B.

Tepix today at 10:55 AM
Sounds like a big improvement.

No mention of weights, just API. When will the updated weights be released?

miyuru today at 7:48 AM
Judging by the openrouter leaderboard ranking for today, it looks like Dv4F us more popular than mimov2.5.

https://openrouter.ai/rankings?view=day#leaderboard-table

These days cost per task is more important, and SOTA models have become expensive.

KronisLV today at 7:13 AM
Wonder how good the proper version of V4 Pro will be.

I'm still considering pulling the trigger on the annual subscription of Kimi for K3 but it's sometimes slower than I'd like (at least when compared to Anthropic) even on their Vivace plan, and the token limits on the GLM Coding subscription for GLM 5.2 were too easy to hit.

HyperL0gi today at 11:42 AM
Is anyone using DSv4 for their agents that is not related to writing code? Curious about use cases specially for someone using gpt-5.4 mini for classification, categorization, etc
vladukha today at 8:24 AM
Where do you guys get deepseek? I'm hearing a lot of good reviews and want to try it with my pi config. from the deeepseek themselves, openrouter, or anywhere else? does it make a difference? [edit]: whoa it is really fast. will take some time to evaluate quality thou
wolttam today at 7:00 AM
Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.
egeozcan today at 7:31 AM
Every time I want to have fun coding something with natural language processing, I use deepseek flash. It's just incredible for the price. I have a fairly popular app with 400 users that uses DeepSeek in the background and it still didn't hit even 50 bucks of usage in a month.
PhilippGille today at 7:37 AM
The previous V4 version wasn't called “Preview” by most inference providers. For example, the OpenRouter model slug was `deepseek/deepseek-v4-flash`. So now there will be confusion when someone talks about V4 Flash or when someone offers V4 Flash inference.

Why not call it V4.1?

w2seraph today at 6:40 PM
This made my day !
flysoft today at 7:33 AM
Finally have a model with usable intelligence, at a reasonable price. Can't imagine what Pro GA would look like, considering pro preview has only 1.6t parameters.
throwaw12 today at 8:57 AM
how different is their harness from Pi coding agent harness, is it possible to make an extension for Pi which can implement deepseek harness?
storywatch today at 7:43 AM
How's their performance in English prose? We are currently searching for cost effective ways to keep story wikis up to date.
nathaah3 today at 8:55 AM
DS v4 flash has been my goto model for tasks in work. its been unsurprisingly fast and cheap.
k__ today at 10:03 AM
I'd take more throughput while everything else stays the same.
kamikazechaser today at 7:41 AM
The flash variant is on par with Sonnet 5 on DeepSWE (54%). Big, if true.
znnajdla today at 10:05 AM
The conspiracy theorist in me wants to think that the 80% drop in GPT 5.6 Luna prices today is correlated with this update from DeepSeek. Perhaps OpenAI has already hacked its competitors with it's Mythos-like models and is aware of what competitors are doing and is able to react in advance.
spwa4 today at 6:27 AM
In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.
XCSme today at 11:51 AM
Can't really use it now, without giving away your data:

> Trains: this provider may use prompts for training and may retain prompt data.

ra today at 8:03 AM
What's the best way to run this on a 64GB M2 Pro?
sparse-Matrix today at 10:35 AM
This may come as a surprise to a lot of AI concerns, but I have -zero- interest in paying for a model.
sreekanth850 today at 12:44 PM
how this compare to luna high with reduced pricing.
mekky16 today at 10:07 AM
if they were anthropic they wouldve just released it as a new model
deleted today at 6:31 AM
truth_seeker today at 9:54 AM
The magic of post training with valuable dataset
sourcecodeplz today at 8:48 AM
i've made a comparison between this and GPT Luna (recent %80 price drop)

https://x.com/SourceCodeplz/status/2083099712760987746

i prefer GPT-5.6 Luna honestly

dakolli today at 10:00 AM
Chinese labs rushing to release models this week, because it's inevitable that Washington regulates Chinese models in the next 4 weeks. All the US AI leaders have been taking trips to Washington this week, what do you think they're there for..
Tepix today at 10:50 AM
[dead]
tosh today at 7:22 AM
[dead]
dnhkng today at 6:16 AM
DeepSeek V4 Flash (Preview → 2026-07-31)

• Terminal Bench: 56.9 → 82.7 (+25.8)

• Toolathlon: 51.8 → 70.3 (+18.5)

Compared to GPT-5.6 Terra:

• Terminal Bench: Flash 82.7 vs Terra 78.4

• Toolathlon: Flash 70.3 vs Terra 53.1

• DeepSWE: Flash 54.4 vs Terra 69.6

• Agents' Last Exam: Flash 25.2 vs Terra 50.4

Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

try-working today at 7:40 AM
Let's see how the market reacts.
Havoc today at 8:40 AM
>benchmark results far exceeding V4-Pro-Preview:

Wow that's crazy

Good times for those that don't need strict data protection