CourionAI
EN
Newsletter
← All news
models 3 min read

DeepSeek's Cheap Model Can Now See, and It Edges Past Opus 4.8 on Two Visual Tests

V4-Flash-Vision-Exp adds image understanding to DeepSeek's smaller model. Images are billed at up to 384 tokens each, at the same price as text.

An eye icon overlaid on a photo frame with a small chip in the corner

DeepSeek released an experimental version of its smaller model on 21 August that can look at pictures. V4-Flash-Vision-Exp is live on the company’s API platform, keeps the text abilities of the existing V4-Flash, and adds image understanding on top.

The two numbers people are quoting come from visual benchmarks. On ALE, a set of long multi-step application tasks, DeepSeek’s model scored 27.3 against Anthropic’s Opus 4.8 at 25.7. On ZeroBench, a collection of 100 deliberately hard image analysis puzzles, it reached 35.0 to Opus 4.8’s 34.0. Both are narrow wins on two tests, not a general claim of superiority, and DeepSeek itself phrases it more modestly: multimodal agent performance is now “close to Opus-4.8”. The base model underneath is a mixture-of-experts design, meaning the full model holds a large pile of parameters but only activates a small slice of them for any given request, which is how it stays cheap to run.

The practical details matter more than the leaderboard. Images are converted into tokens for billing at up to 384 tokens each, charged at the ordinary V4-Flash rate, so adding a screenshot to a request costs roughly what a few paragraphs of text would. DeepSeek also switched on a Files API the same day: you upload an image once, get a file ID back, and reference that ID in later requests instead of re-uploading. It is free, and it saves a lot of bandwidth if you are sending the same document or diagram repeatedly.

The word “experimental” in the model name is doing real work. DeepSeek has used the “-Exp” suffix before for models that get tested in public, adjusted, and later folded into a stable release. Do not build anything you depend on around it yet. Worth knowing too: this is an API-only release. DeepSeek has published open weights for several of its models before, including V4-Flash in July, but nothing has been published for the vision variant so far. The broader pattern is that vision is no longer a premium feature reserved for the biggest and most expensive models. Once a cheap model can reliably read a screenshot, a chart or a scanned invoice, a lot of everyday automation becomes affordable that previously was not.

What this means for you: if you are just curious about AI, the takeaway is that “can it look at my screenshot” is quickly becoming a standard feature rather than a selling point. If you build with these tools, the interesting number is the price, not the benchmark: image input at text rates makes document processing, screenshot debugging and chart reading cheap enough to try on real volume. A fair caveat before you get excited: experimental models can change or disappear, and the benchmark gaps here are around one point, which is well inside the range where a different test would flip the ranking.

Sources

Source: https://api-docs.deepseek.com/news/news260821/

Next story

Google's Free Models Passed a Billion Downloads. Some of Them Are Running in Orbit

Gemma crossed one billion downloads and 100,000 community variants in two years. Alibaba's Qwen claimed three billion five days earlier.

A small rocket orbiting a planet made of stacked download arrows