deepseek-flash

Operational
Ranked #23 in the capability rankings
Reasoning
Tools
Vision
Multimodal
Open weights
1M

Standard & Paid Price

Standard Price
MeterStandard Price
Input price$0.1500 / 1M tokens
Completion price$0.6000 / 1M tokens
Cache read price$0.0030 / 1M tokens
Dynamic pricing · 2 tiers · Time rules
Billing details
Vendor list price comparison

Description

DeepSeek V4.1 Flash — the canonical model id for new integrations. It supports text and image input, thinking and non-thinking modes, tool calling and JSON output. The retired deepseek-v4-flash and deepseek-v4-flash-vision-exp ids remain compatibility aliases served by this same model at the same Flash price. 1M-token context window, 384K maximum output. On /v1/responses, send the whole conversation in the input array on every turn: no conversation state is kept upstream, so requests that carry previous_response_id or conversation are refused.

Capabilities

Input modalities
text, image
Output modalities
text
Tasks
chat, reasoning, vision

Decision evidence

DeepSeek V4.1 Flash, served as deepseek-flash since 2026-09-10, is an MIT-licensed open-weight mixture-of-experts model with a 1M-token context window and text and image input. Reviewed tests find it cheap to run and competitive but uneven from task to task, and DeepSeek names no production engine for self-hosting it.

2026-09-25 · Markdown evidence · JSON data

API endpoints

  • POST /v1/chat/completions (openai)
  • POST /v1/responses (openai-response)
  • POST /v1/messages (anthropic)