DeepSeek-V4-Flash-Vision-Exp: a Vision Model for the Agent Era
DeepSeek's first multimodal model (284B/13B-active MoE, 1M context) takes image input at ≤384 tokens each. Multimodal agent benchmarks approach Opus 4.8 while text ability stays level with V4-Flash — a one-line model swap to upgrade.