Research
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Overview Research area: Natural Language Processing — specifically large language model architecture, long-context inference efficiency,

- arXiv
- 2609.19969
- Published
- 2026-09-17
- Authors
- DeepSeek-AI, :, Anyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian, Bingzheng Xu, Bochao Wu, Bowei Zhang, Boyi Deng, C. C. Yu, Chao Jin, Chaofan Lin, Chen Dong, Chenbing Wang, Chenfan Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyuan Zhang, Chenhao Xu, Chenqi Zhao, Chenze Shao, Chuhao Wang, Chuqi Zhang, Damai Dai, Dejian Yang, Deli Chen, Di Huang, Di Wu, Donghao Li, Erhang Li, Eric Fu, F. Zhou, Fangwei Zhou, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanglin Li, Guanting Chen, Guoai Cao, Guofan Fan, Guolai Meng, Guowei Li, Haichuan Zhang, Haiyang Ma, Haiyang Shen, Han Li, Han Yu, Han Zhang, Hangyuan Deng, Hanwei Xu, Hanxiang Xu, Hanxun Zhong, Hao Guo, Hao Jiang, Hao Li, Hao Qin, Haodong Wen, Haofen Liang, Haofeng Huang, Haohua Liu, Haoling Zhang, Haoming Luo, Haoran Yang, Haotian Xu, Haotian Yuan, Haoting Huang, Haowen Luo, Haoyang Cai, Haoyu Chen, Haozhe Ji, Hengran Zhang, Hengrui Wang, Hengxu Wu, Honghui Ding, Hongxuan Tang, Huadong Wang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J. Yang, J. H. Jin, J. H. Zhang, J. X. Zou, Jia Yu, Jiahui Zhou, Jiajun Chen, Jialiang Huang, Jialin Zhao, Jiamin Tang, Jian Zhou, Jianan Tong, Jianwen Li, Jiaqi Zhu, Jiarui Wang, Jiasheng Ye, Jiashi Li, Jiaxin Xu, Jiaying Ding, Jibai Lu, Jiewen Hu, Jin Yan, Jincheng Zhai, Jingchang Chen, Jingcheng Hu, Jingli Zhou, Jingsheng Xu, Jingting Xiang, Jingyan Yun, Jingyang Yuan, Jingyuan Cheng, Jinhua Zhu, Jinpeng Wang, Jinyi Chen, Jinyi Hu, Jiping Yu, Jueliang Guo, Junbo Pei, Junbo Sun, Junguang Jiang, Junjie Qiu, Junkang Zhou, Junqi Liu, Junren Li, Junxian Li, Junxiao Song, Junyi Guo, Kai Dong, Kaifeng Chen, Kaige Gao, Kang Guan, Kangdong Yuan, Ke Hong, Ke Xu, Kefan Zhao, Kexin Ji, Kexin Zhang, Kexing Zhou, Kuai Yu, Lan Zhang, Lean Wang, Lecong Zhang, Lei Wang, Letian Gao, Liang Zhao, Liansheng Xu, Lihua Guo, Lingxiao Luo, Lingyue Fu, Litao Deng, Litong Wang, Liyue Zhang, Longhao Chen, Lu Chen, Luotian Huang, Luyao Ma, Luyao Wang, M. S. Di, Max Mei, Menghao Ye, Miao Cui, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingjing Zhang, Mingqi Wei, Mingshu Chen, Mingxing Liu, Mingxu Zhou, Mingyu Xu, Mingyu Yang, Mingze Wang, Muyang Chen, Ni Shentu, Ning Wang, Niufang Ning, Panpan Huang, Peixin Cong, Peiyi Wang, Peiyuan Xin, Pengfei Ren, Pengfei Yan, Pengle Zhang, Qi Kang, Qi Tang, Qiancheng Wang, Qiang Li, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Qizhou Guo, Rongxian Xu, Rui Ding, Rui Hu, Rui Tian, Rui Yu, Ruidong Zhu, Ruifan Xu, Ruihan Yang, Ruihang Xia, Ruijie Lu, Ruilin Geng, Ruipeng Hong, Ruiqi Ge, Ruisong Zhang, Ruize Sun, Ruizhe Pan, Runji Wang, Runqian Chen, Runxin Xu, Ruohong Tian, Ruomeng Shen, Ruoyu Zhang, Ryan X., S. H. Liu, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoheng Nie, Shaoyuan Chen, Shengding Hu, Shengkai Lin, Shengwen Ran, Shengyu Liu, Shengyuan Jia, Shi Bai, Shi Feng, Shicheng Xu, Shichun Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shiyuan Feng, Shufan Gong, Shuhan Lin, Shuiping Yu, Shunfeng Zhou, Shuo Yang, Shuomeng Wang, Shuting Guo, Shuting Pan, Shuying Yu, Sinuo Cao, Siyi Lin, Sizhe Chen, Songyang Chen, Songyang Zhou, Tao Ni, Tao Yun, Tian Jin, Tian Pei, Tian Ye, Tianle Lin, Tianran Ji, Tianyi Cui, Tianyuan Yue, Tingting Yu, Tongrui Xiong, Wangding Zeng, Wei Liu, Wei Zhang, Weibin Xu, Weihao Zeng, Weilin Zhao, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjing Yao, Wenjun Gao, Wenkai Shao, Wenkai Yang, Wenli Zhang, Wenlu Wang, Wenlve Huang, Wenqian Yan, Wentao Zhang, Xi Gao, Xiang He, Xiang Li, Xiangli Li, Xiangwen Wang, Xiangying Zhang, Xiankui Wei, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaojian Qu, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xiaoyao Zou, Xiaoyuan Li, Xicheng Guo, Xieting Chu, Xin Cheng, Xin Liu, Xin Xie, Xinbo Xu, Xingchao Liu, Xingchen Liu, Xingkai Yu, Xingyou Li, Xintong Yao, Xinyang Chen, Xinyong Jiang, Xinyu Yang, Xinyu Yang, Xu Chen, Xuanyu Wang, Xubei Zhong, Xuecheng Su, Xuejie Liu, Xuheng Lin, Xujie Fan, Xuncheng Zhao, Xuwei Fu, Y. C. Yan, Y. H. Jiang, Y. T. Wu, Y. W. M., Y. Z. Wang, Yafei Gao, Yang Yang, Yang Zhang, Yanru Ma, Yanwen Huang, Yao Li, Yao Li, Yao Meng, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yaoyang Ye, Yehang Yin, Yexinrui Wu, Yi Qian, Yi Tao, Yi Yu, Yichao Zhang, Yichen Jiang, Yicheng Wang, Yifan Ding, Yifan Shi, Yifeng Peng, Yifeng Zhai, Yijia Wu, Yiliang Xiong, Yilun Wang, Ying He, Ying Zhou, Yingjia Luo, Yinmin Zhong, Yiping Wang, Yisong Wang, Yixiang Zhang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyao Yang, Yiyuan Liu, Yizai Cai, Yizhen Wei, Yizhi Wang, Yonglun Yang, Yongqi Zhuo, Yongqiang Guo, Yongtong Wu, Yu Wu, Yu Zhang, Yuan Bian, Yuan Cheng, Yuan Ou, Yuan Sun, Yuanfan Xu, Yuanhang Sun, Yuanhao Li, Yuchen Liu, Yuchen Yao, Yudong Han, Yuduan Wang, Yuhan Wu, Yuhao Meng, Yuheng Zou, YuKun Li, Yunchuan Wang, Yunfan Xiao, Yunfan Xiong, Yupeng Chen, Yuqian Cao, Yuqian Wang, Yuqing Chen, Yushun Zhang, Yutong Lin, Yuwei Xiao, Yuxian Gu, Yuxiang Chen, Yuxiang Huang, Yuxiang Luo, Yuxiang You, Yuxin Chen, Yuxin Xiang, Yuxuan Liu, Yuxuan Zhou, Yuyang Zhou, Yuzhe Guo, Yuzhen Huang, Yuzhuo Bai, Z. Y. Z., Zanlin Ni, Zehao Wang, Zehua Zhao, Zehui Ren, Zejun Zhao, Zhangli Sha, Zhanying Wang, Zhaochen Zhang, Zhaoshuai Du, Zhe Fu, Zhean Xu, Zhenda Xie, Zheng Liu, Zhengyan Zhang, Zhenhua Dong, Zhewen Hao, Zhibang Wang, Zhibin Gou, Zhicheng Ma, Zhihao Li, Zhihong Shao, Zhihuan Huang, Zhijie Li, Zhirui Lu, Zhixian Huang, Zhixuan Chen, Zhixuan Chen, Zhixuan Pan, Zhiyu Wu, Zhizhou Ren, Zhu He, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihui Gu, Zijia Zhu, Zili Zhang, Zilin Li, Zilong Hou, Zilong Lyu, Ziqiao Wang, Ziwei Xie, Ziya Zhang, Ziyi Gao, Zizheng Pan, Zonglin Li, Zongqing Yao, Zui Chen, Zuofan Wu, Chenchen Ling, Chengyu Hou, Chong Chen, D. Li, Di Qi, Dongjie Ji, Fang Wei, Fanyi Xia, Fei Xie, Feiyi Tan, Hailong Guo, Haiyan Zhai, Hui Zhou, Huihui Tan, Huijie Li, Jia Luo, Jia Song, Jialu Cai, Jian Liang, Jiangting Zhou, Jiaqi Gao, Jiayi Shao, Jie Chen, Jieyu Yang, Jin Chen, Jingde Zhang, Jingzi Zhou, Jinqian Wang, Jinyang Liu, JinZhao Sun, Junhua Ling, Junmin Zheng, Kaicheng Yang, Ke Xu, Le Su, Leyi Xia, Liangfeng Ding, Lin Zhuo, Linwang Ma, Linyan Zhu, Liyu Cai, Luqi Yao, M. K. Zhang, Meng Li, Miao Lin, Miaojun Wang, Min Zhang, Mingming Li, Mingming Wang, Mingze Yin, Minmin Han, Nan Cao, Ning Wang, Ningxin Ma, Panpan Wang, Peihan Lin, Peng Sun, Peng Zhang, Qian Ying, Qiang Xiang, Qiao Wang, Qingmiao Mao, Qiwei Jiang, Rongli Jin, Ruyi Chen, Sha Tao, Shangmian Sun, Shaoqing Wu, Shichao Zou, Si Lei, Tianyang Zhang, Tianyu Sun, Tingting Yin, W. L. Xiao, Wei An, Wei Li, Wei Wang, Weiwei Lin, Wenqing Hou, X. Lin, Xiangfei Meng, Xianzhu Huang, Xiao Peng, Xiaoqian Li, Xiaoting Zhang, Xiaowen Sun, Xiaoxiang Wang, Xiaoyu Ye, Xinrou Zhang, Xinyu Zhang, Xue Cao, Xueyin Chen, Yanan Zhou, Yanhong Xu, Yao Xia, Yao Xu, Yi Shao, Yihong Zhang, Yiling Ma, Ying Tang, Yining Lou, Yiru Chen, Yishi Piao, Yixuan Chen, Yong Xiong, Yuchen Xuan, Yuehan Yang, Yuer Xu, Yukun Zha, Yunxian Ma, Yuping Lin, Yuting Yan, Yutong Xie, Yuwen Sheng, Yuxuan Zhu, Zekai Zhang, Zhe Ju, Zhenzhen Lin, Zheren Gao, Zheyang Sun, Zhigang Yan, Zhongyu Wu, Zi Wang, Zihua Qu, Ziling Yan, Ziyi Wan
AI summary
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionOverview
- Research area: Natural Language Processing — specifically large language model architecture, long-context inference efficiency, and KV cache compression for multimodal Mixture-of-Experts (MoE) models.
- Technical level: Advanced. The paper assumes familiarity with attention mechanisms (sliding-window attention, sparse attention, grouped-query attention, multi-head latent attention), Mixture-of-Experts routing, quantization formats, and inference system engineering.
- Scope: The paper introduces a 552B-parameter multimodal MoE model with a Causal Encoder-Decoder architecture, Compressed Sparse Attention 2, and FP4 KV caching, reporting KV cache footprint reductions of roughly 4x to 437x across DeepSeek model generations plus post-training results on reasoning, agentic, and multimodal benchmarks.
What This Paper Is About
Long-horizon AI agents generate input-heavy workloads, and the resulting KV caches strain HBM capacity, SSD capacity, and data-transfer bandwidth — the paper states these compute, storage, and bandwidth demands are together the primary bottleneck to lowering deployment costs. DeepSeek-V4.1-Flash attacks this problem by compressing the KV cache across model architecture, cache precision, and deployment strategy simultaneously, while attempting to improve model quality rather than trade it away.
Key Contributions
-
A compressed KV cache footprint at the model-architecture level. Compressed Sparse Attention 2 (CSA2) reuses global KV (main KV plus indexer K) and Top-K indices across layers through three statically assigned modes (Full, Reindex, Reuse). Combined with FP4 main KV caching, this reduces the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash's.
-
A deployment-level cache reduction via SWA Bounded Replay. Instead of exactly replaying
L × n_wintokens to reconstruct SWA KV states (whereLis layer count andn_winis the SWA window size), the method replays only the most recentn_wintokens. This reduces the persistent KV cache footprint (on SSD or in host memory) to roughly 1/8 of DeepSeek-V4-Flash's. -
A streamlined and extended architecture. The paper introduces the Causal Encoder-Decoder (CED), Single-Pass mHC with a Mega-mHC deployment kernel that halves activation memory traffic relative to the original four-kernel implementation, the Engram conditional memory module (196B parameters), DSpark speculative decoding, and a Hierarchical Sparse Indexer.
-
A pretrained and post-trained multimodal model. DeepSeek-V4.1-Flash is pretrained on a 45T-token multimodal corpus, with sparse attention trained from scratch at 64K sequence length without dense attention warmup, and post-trained using supervised fine-tuning, reinforcement learning, and on-policy distillation.
Main Findings
- Global KV cache footprint: 890 bytes per token, held in HBM, roughly 1/4 of DeepSeek-V4-Flash's footprint. Figure 1(b) reports approximately 4-fold and 437-fold reductions in per-token global KV cache size relative to DeepSeek-V4-Flash and DeepSeek-V1, respectively.
- Persistent KV cache footprint: reduced to roughly 1/8 of DeepSeek-V4-Flash's through SWA Bounded Replay, with the paper reporting only negligible performance degradation from the approximate reconstruction.
- Parameter activation asymmetry: the model has 552B backbone parameters and 196B Engram parameters, activating 8B parameters per token during prefill and 16B during decode.
- Decode FLOPs stay nearly flat with context: extending context length 256-fold from 4K to 1M increases single-token Decode FLOPs by only 1/4. The accounting weights BF16, FP8, and FP4 operations by 1, 0.5, and 0.25 respectively.
- Base model efficiency: DeepSeek-V4.1-Flash-Base achieves world knowledge, reasoning, and coding abilities comparable to DeepSeek-V4-Pro-Base, with 5%–10% improvements on held-out evaluations, using only 1/3 total parameters and 1/4 activated parameters.
- Reasoning: sustained high accuracy on reasoning-intensive benchmarks such as mathematics and competitive programming, described as comparable performance with top open-source models like Kimi-K3 and DeepSeek-V4-Pro.
- Agentic: performance on par with closed-source frontier models across Terminal-Bench 2.1, DeepSWE v1.1, and AutomationBench, with the paper acknowledging a gap versus giant models on science-oriented agentic tasks such as Terminal-Bench 4.0 that require expert-level domain knowledge.
- Multimodal: surpasses top-tier open-source competitors like Kimi-K3 specifically on benchmarks evaluating visual reasoning and interpretation of professional charts; the paper acknowledges a distinct overall performance gap compared to giant closed-source systems.
- Practical task coverage: the paper states the model is capable of completing over 95% of real-world tasks.
- Kernel efficiency: each CSA2 Reuse Mode layer executes with only 15 kernels during prefill and 11 during decode.
- FP4 format choice: the paper selects E2M1 with one E4M3 scale per 16 channels, following NVFP4 but omitting its second-level global scale. It reports the format supports magnitudes up to 448 × 6 = 2688, that the largest trained RMSNorm weight magnitude is approximately 1, that the L2 norm of the 512-channel KV latent is at most approximately sqrt(512), that the maximum absolute value across channels after RoPE rotation is bounded by approximately sqrt(512) ≈ 22.6, and that the maximum magnitude observed during training is around 10. FP8 is retained for the SWA KV cache due to its sensitivity to quantization.
Note: specific numeric scores for the named benchmarks (Terminal-Bench 2.1, DeepSWE v1.1, AutomationBench, Terminal-Bench 4.0) are not reported in the available paper content.
Methodology in Plain English
The team started from the observation that DeepSeek-V4 decomposes into a sliding-window-attention local backbone plus a compressed global context branch. They chose to simplify the global branch aggressively while keeping the local attention design mostly intact.
Along the model dimension, they built CSA2, which exploits three multiplicative cost dimensions at once: entry size, sequence compression, and layer reuse. Instead of every layer computing and storing its own global KV and its own Top-K index selection, layers are statically labeled Full, Reindex, or Reuse. Full layers do the complete job; Reindex layers borrow a preceding layer's KV and indexer K but rescore with their own query to pick fresh Top-K entries; Reuse layers borrow both the KV and the already-computed Top-K indices and skip indexing entirely. Every layer still computes its own global query and its own sliding-window KV. They also removed the overlapping source entries and absolute positional embeddings that the earlier CSA compressor used, and they derive indexer K by projecting main KV entries rather than through a separate compression path.
Because index reuse alone does not shrink the scanning cost of the indexers that remain, they added a Hierarchical Sparse Indexer in the decoder only. The first Full Mode layer scores the full causally visible range and, in addition to its own Top-K selection, builds a candidate pool by picking blocks with the highest maximum index scores — for example, 2,048 blocks of 8 positions each yields 16,384 candidate positions. Later Reindex layers search only within that pool. For a fixed pool size this makes the per-query cost of deeper indexers constant rather than linear in context length, though the first Full Mode layer still scans the full range.
For prefill cost, they separated the Transformer into a 20-layer causal encoder and a 20-layer decoder, so that decoder global KV is projected from the encoder's final hidden state using layer-dependent projection weights rather than computed from each decoder layer's own hidden state. This halves prefill complexity from O(NL) to approximately O(NL/2). Since sliding-window attention is still computed layer-wise, decoder SWA replay would normally require processing an extra n_win × L/2 tokens; SWA Bounded Replay instead prefills only the last n_win tokens, motivated by prior evidence that SWA's effective receptive field is much smaller than its theoretical one.
On precision, they applied quantization-aware training to the main KV cache during post-training, quantizing after RoPE, and used an FP4 format with per-16-channel scales. They also reworked residual-stream mixing: Single-Pass mHC shifts input-mixing coefficients by one block so each block consumes coefficients produced by the previous block, removing a dependency that previously forced an extra read of the residual stream, and the Mega-mHC kernel fuses residual update, input mixing, and coefficient prediction into one pass. Training used head-wise Muon for Query and Key weights, and a momentum-based update with Sinkhorn balancing for the Engram embedding tables, token embedding, and prediction head.
Post-training deliberately introduced no algorithmic innovation: it followed supervised fine-tuning, then reinforcement learning, then on-policy distillation, with all substantive changes in the data pipeline — large-scale automated pipelines for data synthesis and environment construction, plus progressive scaling of data, tasks, and rollouts.
Why This Matters
The paper reframes the long-context cost problem: prior sparse-attention work reduced computation, but the paper argues persistent storage and data movement have become the more prominent bottlenecks. By reducing KV cache footprint at the architecture, precision, and deployment levels at once, the work targets the constraints that actually limit serving throughput.
Impact on research: The paper argues that none of the prior layer-dimension compression methods (IndexCache, YOIO, HySparse) covers all three multiplicative compression dimensions jointly — index reuse alone saves no main KV storage, network-wide routing sharing limits performance, and hybrid designs still retain full attention layers. CSA2 is presented as combining all three. The SWA Bounded Replay result is framed as establishing a new storage–computation trade-off that other long-context systems could adopt.
Real-world applications (as described or implied by the paper):
- Long-horizon agent deployment where frequent tool calls generate extensive prefill requests and KV cache misses.
- Coding tasks and white-collar workflows, which the paper states the model is fully capable of handling.
- Frontend development and office automation, where the model uses rendered screen captures for visual inspection and self-correction.
- Visual agentic workflows requiring interpretation of professional charts.
Industry relevance: The paper explicitly targets lowering the cost barrier to deploying long-horizon agents at scale, emphasizing that the small activation footprint yields low inference latency and serving cost. It characterizes the model as a fast, affordable assistant supporting daily work for a broad population of users, and positions it as a new starting point for continued scaling of architecture, pre-training, and post-training.
Future Directions
- Closing the science-oriented agentic gap. The paper acknowledges a remaining gap versus giant models on tasks like Terminal-Bench 4.0 that demand expert-level domain knowledge.
- Closing the multimodal gap. The paper states a distinct overall performance gap remains compared to giant closed-source systems, despite surpassing competitors like Kimi-K3 on visual reasoning and professional chart interpretation.
- Joint scaling. The paper states it will pursue joint scaling of model architecture, pre-training, and post-training to further explore the frontier of model intelligence, positioning DeepSeek-V4.1-Flash as a starting point rather than an endpoint.
- Extending the storage–computation trade-off. The SWA Bounded Replay finding is presented as a new trade-off point; the question of how far the replay bound can be pushed before degradation becomes non-negligible is left open by the paper's description of the result as incurring only negligible degradation.
Target Audience
This paper benefits most readers working on large-scale LLM inference systems and model architecture: inference infrastructure engineers managing HBM/SSD capacity and data-transfer bandwidth, researchers working on sparse attention and KV cache compression, and teams building or deploying long-context agentic systems. It is also relevant to practitioners interested in FP4 quantization for KV caches, Mixture-of-Experts training at scale, and multimodal model design. Given the density of architectural detail, quantization reasoning, and kernel-level accounting, readers without a background in transformer internals and inference systems will find it difficult.
Authors’ abstract
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.