The deepseek-v4-flash-vision-exp endpoint adds image input to the V4 Flash line, but the experimental label matters: the cited launch evidence does not establish production reliability, so use a logged pilot with a fallback before considering it a production default.
The API decision in 30 seconds
DeepSeek V4 Flash Vision Exp is a good fit when an existing V4 Flash workflow needs to read screenshots, charts, documents, or other images through an API-compatible interface. For identity-sensitive or high-stakes visual decisions, keep a fallback and validate the task separately.
| Situation | Best input method | Why |
|---|---|---|
| Small local image used once | Base64 data URL | No public hosting required |
| Image already hosted publicly | External URL | Small request payload |
| Large image or repeated reuse | Files API file_id | Reuse the upload and allow up to 64 MiB per referenced image |
| Need to reduce detail for a broad task | detail: "low" | Downscales the image to 512 x 512 before inference |
The exact model string is deepseek-v4-flash-vision-exp. DeepSeek lists the model as experimental and says it is available on the API platform as of August 21, 2026, in its official change log. The release note reports parity with V4 Flash on pure-text capabilities and a large improvement on agent benchmarks that require visual understanding.
Send one image with Chat Completions
The OpenAI-compatible Chat Completions request puts text and an image in a content array inside a user message. The official Vision guide documents the exact model-specific behavior; sending an image to ordinary deepseek-v4-flash returns a 400 error.
import base64
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url="https://api.deepseek.com",
)
with open("chart.png", "rb") as image_file:
encoded = base64.b64encode(image_file.read()).decode("utf-8")
response = client.chat.completions.create(
model="deepseek-v4-flash-vision-exp",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Extract the three trends from this chart."},
{
"type": "image_url",
"image_url": {
"url": f"data:image/png;base64,{encoded}",
"detail": "original",
},
},
],
}
],
)
print(response.choices[0].message.content)
Images are supported in user messages for Chat Completions. Keep the image and the instruction in the same content array so the model receives the visual context and the task together.
Pick the right image transport
Base64 for a small local file
Base64 is the simplest route for a local, one-off image. It avoids public hosting, but the encoded data counts toward the 48 MiB request-body limit and the source image is limited to 32 MiB.
Use it for one-off user or worker uploads, not images reused across a batch.
Public URL for hosted assets
Public http or https URLs keep requests small but must be reachable, under 8,192 characters, downloadable within 60 seconds, and no larger than 32 MiB. Private, expired, or internal URLs can fail before DeepSeek fetches the image.
Files API for reuse or larger files
Upload the image through the Files API, then reference the returned ID in the vision request:
{
"type": "file",
"file_id": "file-api-xxxxxxxxxxxxxxxx"
}
A referenced file can be up to 64 MiB per image and avoids uploading the same bytes for every request. The trade-off is an extra upload and file-lifecycle step; keep the returned ID with the key that created it rather than treating it as a public share link.
The Files API is the practical choice when a file is larger than 32 MiB, the request could exceed 48 MiB, or several agent steps need to inspect the same image.
Control visual detail before you pay for it
The detail field is available for image_url inputs and Responses API image parts; the behavior below follows DeepSeek's official Vision guide.
| Value | Documented behavior | Use it when |
|---|---|---|
low | Downscales to 512 x 512 | Layout, broad scene, or coarse classification is enough |
high | Keeps the original image | Small text or fine details matter |
original | Keeps the original image | You want explicit full-detail handling |
auto | Currently equivalent to original | You accept the current default behavior |
DeepSeek resizes images before inference. The Vision guide says each image has an upper bound of 384 image tokens, and images are counted independently. A very large source image does not necessarily consume proportionally more image tokens after resizing, although large files can still hit upload and request-size limits.
The official Models & Pricing page lists deepseek-v4-flash-vision-exp at the same token rates as V4 Flash: $0.007 per 1M cached-input tokens and $0.22 per 1M cache-miss input tokens during off-peak hours, with peak rates of $0.014 and $0.44. Output is $0.66 off-peak and $1.32 at peak. Image tokens are billed with input tokens, so image count and detail choices still belong in your cost estimate.
Limits that cause real API failures
| Constraint | Limit or behavior |
|---|---|
| Supported formats | JPEG, PNG, GIF, WebP |
| Maximum request body | 48 MiB |
| Maximum Base64 or URL image | 32 MiB |
Maximum Files API file_id image | 64 MiB |
| Maximum images per request | 600 |
Total image size without file_id images | 64 MiB |
Total image size including file_id images | 200 MiB |
| Maximum dimension | 8,192 pixels per side |
| Dimension limit with 15 or more images | 4,096 pixels per side |
| External URL length | 8,192 characters |
| External image download | Must finish within 60 seconds |
Two restrictions are especially easy to miss. Only deepseek-v4-flash-vision-exp accepts images, and image blocks in system or assistant messages fail for Chat Completions. If an image is sent to a non-vision model, DeepSeek documents the 400 error message as This model does not support image.
The same model across three API surfaces
DeepSeek documents the model on three interfaces in its Vision guide:
| Interface | Image block | Result access |
|---|---|---|
| Chat Completions | image_url in a user content array | response.choices[0].message.content |
| Responses API | input_image with input_text | response.output_text |
| Anthropic-compatible API | image at https://api.deepseek.com/anthropic | Anthropic message content |
All three support Base64, public URLs, and Files API references, but the content types differ. Do not copy the Chat Completions block unchanged into the Responses API.
What the launch evidence says, and what it does not
DeepSeek's August 21 change log reports strong launch benchmarks, including Terminal Bench 2.1 at 83.9 and Chartography at 64.3 at p0.95. These are vendor-reported results, not independent reproduction; the release also notes that text-only V4 Flash ignores multimodal elements in two visual evaluations.
The launch benchmark results are vendor-reported, so validate the visual tasks that matter to your application before routing production traffic.
Should you use it in production?
Use DeepSeek V4 Flash Vision Exp for a controlled pilot when your workload is screenshot analysis, chart extraction, document triage, or an agent that needs to inspect visual state. The matching Flash pricing and three input paths make it inexpensive to evaluate, and the 384-token-per-image ceiling gives you a concrete starting point for cost modeling.
Do not make it the sole backend for identity verification, safety decisions, medical interpretation, or other high-consequence visual judgments while the model remains experimental and the cited launch evidence does not establish reliability for those cases. Put a fallback behind the same interface and record the image source, detail setting, input and output usage, latency, retries, and task success.
Before routing production traffic, test at least:
- Small text in screenshots at
lowandoriginaldetail. - Charts with labels, legends, and dense axes.
- Multiple images in one request.
- Private and slow image URLs.
- Tool calls after visual inspection.
- Incorrect or ambiguous identity prompts.
- Fallback behavior after a 400, timeout, or malformed image response.
DeepSeek V4 Flash Vision Exp API FAQ
What is the exact model name?
Use deepseek-v4-flash-vision-exp. DeepSeek's August 21, 2026 change log identifies it as an experimental multimodal model on the API platform.
Is it priced like V4 Flash?
Yes. DeepSeek's pricing page lists the same cache-hit, cache-miss, and output token rates for Vision Exp and V4 Flash. Image tokens are billed with input tokens, with up to 384 image tokens per image after resizing.
Can it generate images?
The official Vision guide documents image understanding, not image generation. Treat this endpoint as understanding-only unless DeepSeek publishes separate generation support.
Why does my request return a 400 error?
Check the model string, message role, content-block type, file size, and image format. Images sent to a non-vision model or placed in unsupported message roles can trigger the documented This model does not support image error.