Mazda CX5
Як зробити цифрову копію себе. Ось відео —>

DeepSeek Releases Multimodal Model V4.1-Flash with Lower Memory Consumption and Record Speed

Chinese AI company DeepSeek has introduced a new model, DeepSeek-V4.1-Flash. The new model is built on the Mixture-of-Experts (MoE) architecture with a total of 552 billion parameters and has native support for visual data analysis. According to the developers, V4.1-Flash is the smallest in the new line of architectures, but surpasses previous flagships in terms of performance and economy.

Leave a comment
DeepSeek Releases Multimodal Model V4.1-Flash with Lower Memory Consumption and Record Speed

Chinese AI company DeepSeek has introduced a new model, DeepSeek-V4.1-Flash. The new model is built on the Mixture-of-Experts (MoE) architecture with a total of 552 billion parameters and has native support for visual data analysis. According to the developers, V4.1-Flash is the smallest in the new line of architectures, but surpasses previous flagships in terms of performance and economy.

Architecture and technical features

By using the new Causal Encoder-Decoder architecture, the model uses only 8 billion active parameters for input and 16 billion for output, Neowin reports . Thanks to new pre-training methods and large-scale reinforcement learning (RL), V4.1-Flash has shown benchmark results that exceed the performance of models such as Kimi-K3, GLM-5.3, Claude Opus 5, GPT 5.6-Sol and its own predecessor DeepSeek-V4-Pro.

The developers have also significantly reduced the load on hardware resources. The model requires 4 times less high-bandwidth memory (HBM) and 8 times less space on SSD drives for storing KV cache compared to the previous generation.

Cache compression can significantly reduce the cost of AI agents, as cache reads account for a significant portion of the total cost of queries.

API changes and deprecation of old models

DeepSeek-V4.1-Flash is now available via the official API. In connection with the release, the company is decommissioning the following legacy models:

  • requests to V4-Flash and V4-Flash-Vision-Exp are automatically redirected to V4.1-Flash;

  • The V4-Pro model is also being phased out due to the complete superiority of V4.1-Flash in speed, price, and quality. Starting September 14 (04:00 UTC), all requests for V4-Pro will be redirected to V4.1-Flash at the new model's rates until the release of the upcoming V4.1-Pro.

V4.1-Flash support has already been integrated by the company's official partners — WorkBuddy_AI and OpenCode, allowing the model to be used for vibe coding immediately after release. Since the model is distributed as open source, it is expected to be integrated by many other developers in the industry in the near future.

DeepSeek taught V4 Flash to “see” images. Why does it require fewer tokens to process them than competing models?
DeepSeek taught V4 Flash to “see” images. Why does it require fewer tokens to process than competing models?
On the topic
DeepSeek taught V4 Flash to “see” images. Why does it require fewer tokens to process than competing models?
DeepSeek introduced the new V4 model and crashed competitors' shares
DeepSeek introduced the new V4 model and crashed competitors' shares
On the topic
DeepSeek introduced the new V4 model and crashed competitors' shares
Read the country's main IT news in our Telegram
Read the country's main IT news in our Telegram
On the topic
Read the country's main IT news in our Telegram

Have important news to share? Message our Telegram bot

Key events and useful links in our Telegram channel

Discussion
No comments yet.