DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
What changed
DeepSeek released DeepSeek-V4.1-Flash, an AI model that can work with up to 1 million units of text or other input at once. It also uses new memory-saving methods intended to make repeated processing of very long inputs more efficient.
What this means for you
Developers building systems for long documents or extended, multi-step tasks may be able to handle larger inputs with less strain on computing resources. The feed does not specify where the model is available or what it costs, so access remains unclear.