Standard & Paid Price
| Meter | Standard Price |
|---|---|
| Input price | $0.1500 / 1M tokens |
| Completion price | $0.6000 / 1M tokens |
| Cache read price | $0.0030 / 1M tokens |
Description
DeepSeek V4.1 Flash — the canonical model id for new integrations. It supports text and image input, thinking and non-thinking modes, tool calling and JSON output. The retired deepseek-v4-flash and deepseek-v4-flash-vision-exp ids remain compatibility aliases served by this same model at the same Flash price. 1M-token context window, 384K maximum output. On /v1/responses, send the whole conversation in the input array on every turn: no conversation state is kept upstream, so requests that carry previous_response_id or conversation are refused.
Capabilities
- Input modalities
- text, image
- Output modalities
- text
- Tasks
- chat, reasoning, vision
Decision evidence
DeepSeek V4.1 Flash, served as deepseek-flash since 2026-09-10, is an MIT-licensed open-weight mixture-of-experts model with a 1M-token context window and text and image input. Reviewed tests find it cheap to run and competitive but uneven from task to task, and DeepSeek names no production engine for self-hosting it.
2026-09-25 · Markdown evidence · JSON data
API endpoints
POST /v1/chat/completions(openai)POST /v1/responses(openai-response)POST /v1/messages(anthropic)