DeepSeek has launched an experimental multimodal version of its V4-Flash model, giving developers access to a system that can process both text and visual inputs. Called DeepSeek-V4-Flash-Vision-Exp, the model is now available through DeepSeek’s API Platform and is designed to handle tasks involving images and screenshots alongside conventional text-based reasoning.

The release marks another step in DeepSeek’s effort to compete with leading AI developers on increasingly capable agentic and multimodal systems. DeepSeek says V4-Flash-Vision-Exp retains the text capabilities of V4-Flash while making a significant improvement on multimodal agent benchmarks, with performance approaching Anthropic’s Opus 4.8 on tests designed to measure how AI systems understand visual information and take actions based on it.

DeepSeek Adds Vision Capabilities To V4-Flash

DeepSeek-V4-Flash-Vision-Exp is an experimental extension of the company’s V4-Flash model. Its main addition is the ability to understand visual prompts, allowing developers to send images and screenshots along with text.

This makes the model suitable for workflows where an AI agent needs to interpret what is displayed on a screen rather than relying solely on structured text or application programming interfaces.

DeepSeek says the experimental model maintains V4-Flash’s capabilities in areas including agents, reasoning and world knowledge. The addition of vision capabilities is intended to expand those capabilities into multimodal workflows.

DeepSeek V4-Flash-Vision-Exp At A Glance

FeatureDetails
ModelDeepSeek-V4-Flash-Vision-Exp
Base modelDeepSeek V4-Flash
Model typeMultimodal
Visual inputsImages and screenshots
Text capabilitiesAgents, reasoning and world knowledge
AvailabilityDeepSeek API Platform
StatusExperimental
Image tokenizationUp to 384 tokens per image
Input methodsBase64, external URLs, Files API
PricingV4-Flash pricing

The model’s availability through an API means developers can incorporate its visual reasoning capabilities directly into their own applications and AI agents rather than using it only through a consumer-facing chatbot.

Multimodal Agent Performance Is The Main Focus

DeepSeek is positioning the release around multimodal agents rather than simply image recognition.

A conventional vision-language model may be able to describe an image, read text from a screenshot or answer questions about a photograph. An agentic multimodal model goes further by using visual information as part of a sequence of decisions and actions.

For example, a model could potentially inspect a software interface, understand what is displayed and use that information as part of a larger task.

DeepSeek says V4-Flash-Vision-Exp makes a major improvement over the text-only V4-Flash model on multimodal agent benchmarks. The company says its performance approaches Anthropic’s Opus 4.8 on those evaluations.

V4-Flash Vs V4-Flash-Vision-Exp

CapabilityV4-FlashV4-Flash-Vision-Exp
Text reasoningYesYes
Agentic tasksYesYes
World knowledgeYesYes
Image understandingLimited to text-based workflowsYes
Screenshot analysisNo dedicated vision capabilityYes
Multimodal agentsLimitedMajor improvement
API availabilityYesYes
Experimental statusNoYes

The key difference is therefore not a replacement of the underlying text model, but the addition of visual understanding that allows the model to operate across multiple types of input.

Images Are Tokenized For API Billing

DeepSeek has also provided a specific mechanism for handling images within the API.

Images are tokenized for billing at up to 384 tokens per image, with the model using V4-Flash pricing. This allows developers to incorporate visual inputs without moving to a separate pricing structure for the experimental model.

The actual cost of a multimodal application will depend on how many images it processes, the amount of accompanying text and the number of model calls required to complete an agentic task.

Image Input And API Details

API ComponentDeepSeek V4-Flash-Vision-Exp
Image tokenizationUp to 384 tokens per image
Pricing basisV4-Flash pricing
Image inputSupported
Text inputSupported
Base64 imagesSupported
External image URLsSupported
Files APISupported
DeploymentDeepSeek API Platform

Supporting multiple image-input methods gives developers flexibility when building applications around the model. A developer can provide visual information through encoded images, remotely hosted content or the platform’s file-handling infrastructure.

Screenshots Open New Agentic Use Cases

Screenshot understanding is particularly important for computer-use agents.

Traditional software agents often interact with applications through structured APIs or browser automation tools. But many real-world workflows still require an AI system to understand what a human sees on a screen.

A multimodal model can potentially identify buttons, menus, forms, charts, documents and other interface elements from screenshots.

This could make vision-based agents useful for software testing, customer support, document processing and other tasks involving interfaces that do not expose convenient machine-readable APIs.

Potential Applications

Use CaseHow Vision Could Help
Software testingInterpret application screens and identify problems
Computer-use agentsUnderstand interfaces visually
Customer supportAnalyze screenshots submitted by users
Document processingInterpret scanned documents
Data analysisRead charts and visual reports
Web automationUnderstand page layouts and visual elements
Enterprise workflowsCombine screenshots with text instructions

These are potential applications based on the model’s multimodal capabilities; DeepSeek’s announcement does not establish that every listed workflow is specifically optimized or validated by the company.

DeepSeek Targets Anthropic’s Agentic Lead

The comparison with Anthropic is significant because Claude models have developed a strong reputation for coding and agentic workloads.

DeepSeek says V4-Flash-Vision-Exp brings its multimodal agent performance close to Anthropic’s Opus 4.8 on relevant benchmarks.

The comparison reflects a broader change in AI competition. Model developers are increasingly competing on an agent’s ability to perform complete tasks rather than simply answer questions.

Frontier AI Competition Is Expanding

CapabilityImportance
Text reasoningCore problem-solving
CodingSoftware development and automation
VisionUnderstanding images and screens
Tool useConnecting models to external systems
Agentic reasoningCompleting multi-step tasks
Cost efficiencyDetermines commercial scalability

DeepSeek’s release therefore fits into a larger industry shift toward models that can combine reasoning, perception and action.

DeepSeek Continues Its Low-Cost Strategy

The pricing structure is another important element of the launch.

The model uses V4-Flash pricing, while images are tokenized at up to 384 tokens each. This could make multimodal experimentation more accessible to developers compared with systems where advanced vision capabilities carry substantially higher costs.

DeepSeek has already become known for putting pressure on the economics of frontier AI by offering models that aim to deliver strong performance at comparatively low prices.

The company’s strategy is particularly relevant to developers building high-volume AI applications, where even small differences in per-token costs can become significant when multiplied across millions of API requests.

What Matters For Developers

FactorSignificance
API accessEasy integration into applications
V4-Flash pricingKeeps costs tied to existing model economics
Image tokenizationProvides predictable billing framework
Multiple input methodsSimplifies application integration
Multimodal reasoningExpands potential use cases
Experimental releaseIndicates capabilities may continue changing

The experimental status also means developers should expect the model and its performance characteristics to evolve.

China And The US Remain In A Tight AI Race

DeepSeek’s latest release arrives as Chinese and US AI companies continue competing at the frontier of model development.

Chinese developers have increasingly focused on producing models that combine strong performance with lower operating costs. Meanwhile, US companies such as Anthropic and OpenAI continue to push capabilities in reasoning, coding, multimodal understanding and AI agents.

DeepSeek’s V4-Flash-Vision-Exp adds another dimension to that competition by targeting multimodal agents.

The broader significance is that the performance gap between leading AI systems is increasingly being measured across specialized tasks rather than by a single general-purpose benchmark.

Multimodal AI Could Become The Default For Agents

The release also points toward a future in which AI agents are expected to understand more than text.

Human workers routinely use visual information when interacting with computers. They read dashboards, inspect documents, interpret charts, recognize interface elements and look at images before deciding what to do next.

For AI agents to perform similar tasks independently, visual understanding becomes increasingly important.

DeepSeek’s V4-Flash-Vision-Exp is therefore part of a wider transition from language models toward systems capable of processing multiple forms of information and using that information to complete tasks.

The Bigger Picture

DeepSeek’s V4-Flash-Vision-Exp release shows how quickly multimodal capabilities are moving from a specialized feature toward a core requirement for advanced AI agents. By adding image and screenshot understanding to V4-Flash while retaining its text-based reasoning and agent capabilities, DeepSeek is targeting applications where AI must interpret both language and visual information.

The release is also significant because DeepSeek is competing on both capability and cost. The model is available through the API Platform at V4-Flash pricing, with images tokenized at up to 384 tokens each. If the reported benchmark performance holds up in real-world applications, the combination could make multimodal agents more accessible to developers and increase pressure on competing AI providers to improve both performance and pricing.

Looking Ahead

The most important test for V4-Flash-Vision-Exp will be how well its benchmark results translate into practical agentic workflows. Developers will need to evaluate how accurately the model interprets screenshots, how reliably it performs multi-step tasks and how much visual input affects overall API costs. Because the model is experimental, its capabilities and production suitability may change as DeepSeek gathers more usage data.

For the wider AI industry, the release reinforces the idea that future agents will need to see as well as reason. The competition is moving toward systems that can understand text, images and interfaces, use tools and complete tasks with less human intervention. DeepSeek’s latest model puts another low-cost contender into that race and could further accelerate the development of multimodal AI agents.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.