DeepSeek-v4-flash-vision-exp

162 points - today at 10:33 AM

Source

Comments

ciberado today at 11:08 AM
DS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising.

> Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens.

> Before inference, every image is automatically resized:

> - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio.

> - Larger images are scaled down while preserving their aspect ratio so that the total pixel count after resizing is roughly that of an 800×800 image.

> As a result, there is an upper bound of 384 tokens per image: for example, a 2000×2000 image and a 5000×5000 image consume the same number of tokens after resizing. When a request contains multiple images, each image is counted independently under the same rule—there is no separate calculation for multi-image requests.

400 tokens per image results in 2,500 images per dollar, if I’m not mistaken.

edit: format.

zmmmmm today at 11:06 AM
> Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image.

It's useful but for OCR and a lot of other applications it needs to be a bit higher (eg: putting in a full A4 / Letter sized page)

pu_pe today at 12:06 PM
5kyn3t today at 12:23 PM
For what do you guys use vision in those models? surveillance is the obvious use case... but are there some "nicer" ways to use it?
gozucito today at 11:29 AM
800x800 is 640,000 pixels, or 0.64 Megapixels. That is less than the resolution of computer screens from 1995, Super VGA which has around 0.79 MPs.

This is useful for a reasonable amount of use-cases, but I think the watershed rez will be around triple that, ~1080p, which is enough for almost anything, except small text and subtle details.

try-working today at 11:59 AM
I main V4 Pro at work now, and at home I route between Pro and Flash based on task. Switched to Opus 4.6 for some tasks at work because I needed image input - horrible. So nice to get image input with DS.

Edit: I see it has limited resolution. Luckily I just built a vision worker plugin for DSH that routes image input to Kimi K2.6 on Cloudflare.

LorenDB today at 11:00 AM
I've heard that DeepSeek v4 Flash 0731 has frequently assumed that it has vision capabilities and then resorts to inventing text-based image analysis tools when it finds that it actually can't see. In that case, this is a great upgrade for the model.

Anecdotally, I had to tell 0731 to refrain from viewing screenshots since it kept breaking its sessions by trying to read images.

BrucecarlL today at 11:31 AM
Congratulations! DeepSeek has finally gained eyes — the dark days are about to be behind us.
v9v today at 11:23 AM
Interesting. Wasn't Deepseek's founder saying that they had explicitly decided not to focus on multimodal models at all and were going text-only because they believed it was enough to achieve AGI?
erikkri today at 12:07 PM
Hello Ox Alpha?
dsrtslnd23 today at 11:17 AM
will this be open weights?
locitra today at 12:17 PM
[flagged]
MagicMoonlight today at 12:30 PM
[dead]
lzy today at 11:15 AM
[dead]
jaksdbvqi37u today at 11:40 AM
[dead]