Kihagyás

Towards Highly Efficient Million-Token Context Intelligence DeepSeek-AI research@deepseek.com Abstract We present a preview version of DeepSeek-V4 series, including two strong Mixture-of- Experts(MoE)languagemodels—DeepSeek-V4-Prowith1.6Tparameters(49Bactivated)and DeepSeek-V4-Flashwith284Bparameters(13Bactivated)—bothsupportingacontextlengthof onemilliontokens. DeepSeek-V4seriesincorporateseveralkeyupgradesinarchitectureandop- timization: (1)ahybridattentionarchitecturethatcombinesCompressedSparseAttention(CSA) andHeavilyCompressedAttention(HCA)toimprovelong-contextefficiency;(2)Manifold- Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train bothmodelsonmorethan32Tdiverseandhigh-qualitytokens,followedbyacomprehensive post-trainingpipelinethatunlocksandfurtherenhancestheircapabilities. DeepSeek-V4-Pro- Max,themaximumreasoningeffortmodeofDeepSeek-V4-Pro,redefinesthestate-of-the-artfor openmodels,outperformingitspredecessorsincoretasks. Meanwhile,DeepSeek-V4seriesare highlyefficientinlong-contextscenarios. Intheone-million-tokencontextsetting,DeepSeek- V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared withDeepSeek-V3.2. Thisenablesustoroutinelysupportone-million-tokencontexts,thereby makinglong-horizontasksandfurthertest-timescalingmorefeasible. Themodelcheckpoints areavailableathttps://huggingface.co/collections/deepseek-ai/deepseek-v4. 100 80 60 40 20 0 SimpleQA HLE Apex Codeforces SWE Terminal Toolathlon Verified (Pass@1) Shortlist (Rating) Verified Bench 2.0 (Pass@1) (Pass@1) (Pass@1) (Resolved) (Acc) )%( 1@ssaP / ycaruccA 1.2 1.0 DeepSeek-V4-Pro-Max Claude-Opus-4.6-Max GPT-5.4-xHigh Gemini-3.1-Pro-High 0.8 90.2 85.9 89.1 32063168 3052 0.6 80.680.880.6 75.6 78.1 75.1 0.4 67.9 68.5 65.4 0.2 57.9 51.8 54.6 0.0 46.245.3 44.4 47.2 48.8 0 25 T 6 oken Po 51 s 2 ition (K) 768 1024 40.039.8 37.7 Knowledge & Reasoning Agentic Capabilities )T( sPOLF nekoT-elgniS DeepSeek-V3.2 DeepSeek-V4-Pro DeepSeek-V4-Flash 3.7× lower 9.8× lower 50 40 30 20 10 0 0 256 512 768 1024 Sequence Length (K) )BG( ehcaC VK detalumuccA DeepSeek-V3.2 DeepSeek-V4-Pro DeepSeek-V4-Flash 9.5× smaller 13.7× smaller Figure1 | Left: benchmarkperformanceofDeepSeek-V4-Pro-Maxanditscounterparts. Right: inferenceFLOPsandKVcachesizeofDeepSeek-V4seriesandDeepSeek-V3.2.


Contents 1 Introduction 4 2 Architecture 6 2.1 DesignsInheritedfromDeepSeek-V3. . . . . . . . . . . . . . . . . . . . . . . . . . 7 2.2 Manifold-ConstrainedHyper-Connections . . . . . . . . . . . . . . . . . . . . . . 7 2.3 HybridAttentionwithCSAandHCA . . . . . . . . . . . . . . . . . . . . . . . . . 9 2.3.1 CompressedSparseAttention . . . . . . . . . . . . . . . . . . . . . . . . . . 9 2.3.2 HeavilyCompressedAttention . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.3.3 OtherDetails . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.3.4 EfficiencyDiscussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.4 MuonOptimizer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 3 GeneralInfrastructures 15 3.1 Fine-GrainedCommunication-ComputationOverlapinExpertParallelism . . . . 15 3.2 FlexibleandEfficientKernelDevelopmentwithTileLang . . . . . . . . . . . . . . 16 3.3 High-PerformanceBatch-InvariantandDeterministicKernelLibraries . . . . . . 18 3.4 FP4Quantization-AwareTraining. . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.5 TrainingFramework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 3.5.1 EfficientImplementationofMuon . . . . . . . . . . . . . . . . . . . . . . . 20 3.5.2 Cost-EffectiveandMemory-EfficientImplementationofmHC . . . . . . . 21 3.5.3 ContextualParallelismforLong-ContextAttention . . . . . . . . . . . . . 21 3.5.4 ExtendedAutomaticDifferentiationforFlexibleActivationCheckpointing 21 3.6 InferenceFramework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 3.6.1 KVCacheStructureandManagement . . . . . . . . . . . . . . . . . . . . . 22 3.6.2 On-DiskKVCacheStorage . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 4 Pre-Training 24 4.1 DataConstruction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 4.2 Pre-TrainingSetups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 4.2.1 ModelSetups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 4.2.2 TrainingSetups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 4.2.3 MitigatingTrainingInstability . . . . . . . . . . . . . . . . . . . . . . . . . 26 4.3 Evaluations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 4.3.1 EvaluationBenchmarks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 4.3.2 EvaluationResults . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 2


5 Post-Training 29 5.1 Post-TrainingPipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 5.1.1 SpecialistTraining . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 5.1.2 On-PolicyDistillation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 5.2 RLandOPDInfrastructures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 5.2.1 FP4QuantizationIntegration . . . . . . . . . . . . . . . . . . . . . . . . . . 34 5.2.2 EfficientTeacherSchedulingforFull-VocabularyOPD . . . . . . . . . . . 34 5.2.3 PreemptibleandFault-TolerantRolloutService . . . . . . . . . . . . . . . 34 5.2.4 ScalingRLFrameworkforMillion-TokenContext . . . . . . . . . . . . . . 35 5.2.5 SandboxInfrastructureforAgenticAI . . . . . . . . . . . . . . . . . . . . . 35 5.3 StandardBenchmarkEvaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 5.3.1 EvaluationSetup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 5.3.2 EvaluationResults . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 5.4 PerformanceonReal-WorldTasks . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 5.4.1 ChineseWriting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 5.4.2 Search . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 5.4.3 White-CollarTask . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 5.4.4 CodeAgent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 6 Conclusion,Limitations,andFutureDirections 44 A AuthorListandAcknowledgment 54 A.1 AuthorList . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54 A.2 Acknowledgment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 B EvaluationDetails 55 3


  1. Introduction The emergence of reasoning models (DeepSeek-AI, 2025; OpenAI, 2024c) has established a newparadigmoftest-timescaling,drivingsubstantialperformancegainsforLargeLanguage Models(LLMs). However,thisscalingparadigmisfundamentallyconstrainedbythequadratic computational complexity of the vanilla attention mechanism (Vaswani et al., 2017), which createsaprohibitivebottleneckforultra-longcontextsandreasoningprocesses. Concurrently, the emergence of long-horizon scenarios and tasks — from complex agentic workflows to massive cross-document analysis — has also made efficient support for ultra-long contexts critical for future progress. While recent open-source efforts (Bai et al., 2025a; DeepSeek-AI, 2024;MiniMax,2025;Qwen,2025)haveadvancedgeneralcapabilities,thiscorearchitectural inefficiencyinhandlingultra-longsequencesremainsakeyimpediment,limitingfurthergains fromtest-timescalingandhinderingfurtherexplorationintolong-horizonscenariosandtasks. Inordertobreaktheefficiencybarrierinultra-longcontexts,wedeveloptheDeepSeek-V4 series,includingthepreviewversionsofDeepSeek-V4-Prowith1.6Tparameters(49Bactivated) andDeepSeek-V4-Flashwith284Bparameters(13Bactivated). Througharchitecturalinnova- tions,DeepSeek-V4seriesachieveadramaticleapincomputationalefficiencyforprocessing ultra-longsequences. Thisbreakthroughenablesefficientsupportforacontextlengthofone milliontokens,usheringinaneweraofmillion-lengthcontextsfornext-generationLLMs. We believe our capability to efficiently handle ultra-long sequences unlocks the next frontier of test-timescaling,pavesthewayfordeeperresearchintolong-horizontasks,andestablishesa necessaryfoundationforexploringfutureparadigmslikeonlinelearning. Compared with the DeepSeek-V3 architecture (DeepSeek-AI, 2024), DeepSeek-V4 series retaintheDeepSeekMoEframework(Daietal.,2024)andMulti-TokenPrediction(MTP)strategy, whileintroducingseveralkeyinnovationsinarchitectureandoptimization. Toenhancelong- context efficiency, we design a hybrid attention mechanism combining Compressed Sparse Attention(CSA)andHeavilyCompressedAttention(HCA).CSAcompressestheKVcaches alongthesequencedimensionandthenperformsDeepSeekSparseAttention(DSA)(DeepSeek- AI, 2025), whereas HCA applies more aggressive compression to the KV caches but keeps dense attention. To strengthen modeling capability, we incorporate Manifold-Constrained Hyper-Connections (mHC) (Xie et al., 2026) that upgrade conventional residual connections. Additionally, we introduce the Muon (Jordan et al., 2024; Liu et al., 2025) optimizer to the trainingofDeepSeek-V4series,leadingtofasterconvergenceandimprovedtrainingstability. ToenableefficienttrainingandinferenceforDeepSeek-V4seriesaswellasproductivede- velopment,weintroduceseveralinfrastructureoptimizations. First,wedesignandimplement asinglefusedkernelforMoEmodulesthatfullyoverlapscomputation,communication,and memoryaccess. Second,weemployTileLang(Wangetal.,2026),aDomain-SpecificLanguage (DSL)tobalancedevelopmentproductivityandruntimeefficiency. Third,weprovideefficient batch-invariantanddeterministickernellibrariestoensurebitwisereproducibilityacrosstrain- ing and inference. Fourth, we incorporate FP4 quantization-aware training for MoE expert weightsandtheindexerQKpathtoreducememoryandcomputation. Fifth,forthetraining framework,weextendtheautogradframeworkwithtensor-levelcheckpointingforfine-grained recomputationcontrol;andweenhancetrainingefficiencywithahybridZeROstrategyforthe Muonoptimizer,cost-effectivemHCimplementationsviarecomputationandfusedkernels,and two-stage contextual parallelism to manage compressed attention. Finally, for the inference framework,wedesignaheterogeneousKVcachestructurewithon-diskstoragestrategiesto enableefficientshared-prefixreuse. 4

ByemployinghybridCSAandHCA,alongwithprecisionoptimizationsoncomputation andstorage,DeepSeek-V4seriesachievesignificantlylowerinferenceFLOPsandasubstantially reducedKVcachesizecomparedwithDeepSeek-V3.2,especiallyinlong-contextsettings. The rightpartofFigure1demonstratestheestimatedsingle-tokeninferenceFLOPsandaccumulated KVcachesizeofDeepSeek-V3.2andDeepSeek-V4series. Inthescenarioof1M-tokencontext, evenDeepSeek-V4-Pro,whichhasalargernumberofactivatedparameters,attainsonly27% of the single-token FLOPs (measured in equivalent FP8 FLOPs) and 10% of the KV cache sizerelativetoDeepSeek-V3.2. Furthermore,DeepSeek-V4-Flash,withitssmallernumberof activatedparameters,pushesefficiencyevenfurther: inthe1M-tokencontextsetting,itachieves only10%ofthesingle-tokenFLOPsand7%oftheKVcachesizecomparedwithDeepSeek-V3.2. Additionally,forDeepSeek-V4series,theroutedexpertparametersutilizeFP4precision. While the peak FLOPs for FP4 × FP8 operations are currently the same as FP8 × FP8 on existing hardware,theycantheoreticallybeimplementedtobe1/3moreefficientonfuturehardware, whichwillfurtherenhancetheefficiencyofDeepSeek-V4series. Duringpre-training,wetrainDeepSeek-V4-Flashon32TtokensandDeepSeek-V4-Proon33T tokens,respectively. Afterpre-training,thesetwomodelscannativelyandefficientlysupport 1M-length contexts. In our internal evaluations, DeepSeek-V4-Flash-Base already surpasses DeepSeek-V3.2-Baseacrossamajorityofbenchmarkswithitsmoreparameter-efficientdesign. DeepSeek-V4-Pro-Basefurtherextendsthisadvantagetosetanewperformancestandardamong DeepSeekfoundationmodels,achievingcomprehensivesuperiorityacrossreasoning,coding, long-context,andworldknowledgetasks. Thepost-trainingpipelineofDeepSeek-V4seriesfeaturesatwo-stageparadigm: theinde- pendentcultivationofdomain-specificexperts,followedbyunifiedmodelconsolidationvia on-policydistillation(LuandLab,2025). Initially,foreachtargetdomain—suchasmathematics, coding,agent,andinstructionfollowing—aseparateexpertmodelistrainedindependently. ThebasemodelfirstundergoesSupervisedFine-Tuning(SFT)onhigh-quality,domain-specific datatoestablishfoundationalcapabilities. Subsequently,ReinforcementLearning(RL)isap- pliedusingGroupRelativePolicyOptimization(GRPO)(DeepSeek-AI,2025),whichfurther optimizesthemodelfordomain-alignedbehaviorsguidedbyrewardmodelstailoredtospecific success criteria. This phase yields a diverse set of specialized experts, each excelling in its respectivefield. Finally,tointegratethesedistinctproficiencies,asingleunifiedmodelistrained throughon-policydistillation,whereintheunifiedmodelactsasthestudentlearningtooptimize thereverseKLlosswithteachermodels. SummaryofCoreEvaluationResults • Knowledge: Inassessmentsofbroadworldknowledge,DeepSeek-V4-Pro-Max,themaxi- mumreasoningeffortmodeofDeepSeek-V4-Pro,significantlyoutperformsleadingopen- sourcemodelsontheSimpleQA(OpenAI,2024d)andChinese-SimpleQA(Heetal.,2024) benchmarks. Regardingeducationalknowledge—evaluatedviaMMLU-Pro(Wangetal., 2024b),HLE(Phanetal.,2025),andGPQA(Reinetal.,2023)—DeepSeek-V4-Pro-Max shows a marginal lead over its open-source counterparts. DeepSeek-V4-Pro-Max has significantlyclosedthegapwiththeleadingproprietarymodel,Gemini-3.1-Pro,despite stilltrailingitintheseknowledge-basedevaluations. • Reasoning: Throughtheexpansionofreasoningtokens,DeepSeek-V4-Pro-Maxdemon- stratessuperiorperformancerelativetoGPT-5.2andGemini-3.0-Proonstandardreasoning benchmarks. Nevertheless,itsperformancefallsmarginallyshortofGPT-5.4andGemini- 3.1-Pro,suggestingadevelopmentaltrajectorythattrailsstate-of-the-artfrontiermodelsby approximately3to6months. Furthermore,DeepSeek-V4-Flash-Maxachievescomparable 5


MTP Modules MTP Loss Prediction Head LM Loss Transformer Block × 𝐿𝐿 Post-Block Mixing DeepSeekMoE Residual Mixing Pre-Block Mixing Post-Block Mixing CSA / HCA Residual Mixing Pre-Block Mixing Embedding Input Tokens Figure2 | OverallarchitectureofDeepSeek-V4series. WeusehybridCSA(CompressedSparse Attention)andHCA(HeavilyCompressedAttention)forattentionlayers,DeepSeekMoEfor feed-forwardlayers,andstrengthenconventionalresidualconnectionswithmHC. performancetoGPT-5.2andGemini-3.0-Pro,establishingitselfasahighlycost-effective architectureforcomplexreasoningtasks. • Agent: Onpublicbenchmarks,DeepSeek-V4-Pro-Maxisonparwithleadingopen-source models,suchasKimi-K2.6andGLM-5.1,butslightlyworsethanfrontierclosedmodels. In our internal evaluation, DeepSeek-V4-Pro-Max outperforms Claude Sonnet 4.5 and approachesthelevelofOpus4.5. • Long-Context: DeepSeek-V4-Pro-Maxdeliversstrongresultsonsyntheticandrealuse caseswitha1-million-tokencontextwindow,surpassingevenGemini-3.1-Proonacademic benchmarks. • DeepSeek-V4-Prov.s. DeepSeek-V4-Flash: DeepSeek-V4-Flash-Maxexhibitslowerper- formance in knowledge evaluations due to its smaller parameter scale. However, it achieves comparable results on reasoning tasks when allocated a larger thinking bud- get. In agent evaluations, while DeepSeek-V4-Flash-Max matches the performance of DeepSeek-V4-Pro-Maxonseveralbenchmarks,itstilltrailsitslargercounterpartonmore complex,high-difficultytasks. 2. Architecture Overall,DeepSeek-V4seriesretaintheTransformer(Vaswanietal.,2017)architectureandMulti- TokenPrediction(MTP)modules(DeepSeek-AI,2024;Gloeckleetal.,2024),whileintroducing several key upgrades over DeepSeek-V3: (1) firstly, we introduce the Manifold-Constrained Hyper-Connections(mHC)(Xieetal.,2026)tostrengthenconventionalresidualconnections; 6


(2)secondly,wedesignahybridattentionarchitecture,whichgreatlyimproveslong-context efficiencythroughCompressedSparseAttentionandHeavilyCompressedAttention. (3)thirdly, we employ Muon (Jordan et al., 2024; Liu et al., 2025) as the optimizer. For the Mixture-of- Experts(MoE)components,westilladopttheDeepSeekMoE(Daietal.,2024)architecture,with onlyminoradjustmentsfromDeepSeek-V3. TheMulti-TokenPrediction(MTP)(DeepSeek-AI, 2024; Gloeckle et al., 2024; Li et al., 2024; Qi et al., 2020) configuration remains identical to thatofDeepSeek-V3. AllotherunspecifieddetailsfollowthesettingsestablishedinDeepSeek- V3(DeepSeek-AI,2024). Figure2illustratestheoverallarchitectureofDeepSeek-V4,andthe detailsaredescribedbelow. 2.1. DesignsInheritedfromDeepSeek-V3 Mixture-of-Experts. AspreviousDeepSeek-seriesmodels(DeepSeek-AI,2024;DeepSeek-AI, 2024),DeepSeek-V4seriesalsoadopttheDeepSeekMoEparadigm(Daietal.,2024)forFeed- ForwardNetworks(FFNs),whichsetsfine-grainedroutedexpertsandsharedexperts. Different fromDeepSeek-V3,wechangetheactivationfunctionthatcomputestheaffinityscoresfrom Sigmoid(·) into Sqrt(Softplus(·)). For load balancing, we also employ the auxiliary-loss-free strategy(DeepSeek-AI,2024;Wangetal.,2024a),augmentedbyaslightsequence-wisebalance lossthatpreventsextremeimbalancewithinindividualsequences. ForDeepSeek-V4,weremove the constraint on the number of routing target nodes, and carefully redesign the parallelism strategytomaintaintrainingefficiency. Furthermore,comparedwithDeepSeek-V3,wereplace the dense FFN layers in the initial several Transformer blocks with MoE layers that employ Hashrouting(Rolleretal.,2021). TheHashroutingstrategydeterminesthetargetexpertsof eachtokenaccordingtoapredefinedhashfunctionwithregardtotheinputtokenID. Multi-Token Prediction. As DeepSeek-V3, DeepSeek-V4 series also set MTP modules and objectives. GiventhattheMTPstrategyhasbeenvalidatedinDeepSeek-V3,weadoptthesame strategyforDeepSeek-V4serieswithoutmodification. 2.2. Manifold-ConstrainedHyper-Connections AsshowninFigure2,DeepSeek-V4seriesincorporateManifold-ConstrainedHyper-Connections (mHC)(Xieetal.,2026)tostrengthentheconventionalresidualconnectionsbetweenadjacent Transformerblocks. ComparedwithnaiveHyper-Connections(HC)(Zhuetal.,2025),thecore ideaofmHCistoconstraintheresidualmappingontoaspecificmanifold,andthusenhancethe stabilityofsignalpropagationacrosslayerswhilepreservingmodelexpressivity. Thissubsection brieflyintroducesthestandardHCanddescribeshowwedesignmHCforstabletraining. Standard Hyper-Connections. The standard HC expands the width of the residual stream byafactorof𝑛 hc . Specifically,theshapeoftheresidualstreamisexpandedfrom R𝑑 to R𝑛 hc ×𝑑 , where 𝑑 is the hidden size of the actual layer input. Let 𝑋 𝑙 = [x𝑙,1 ;...;x𝑙,𝑛 hc ]𝑇 ∈ R𝑛 hc ×𝑑 be the residual state before the 𝑙-th layer. HC introduces three linear mappings: an input mapping 𝐴 𝑙 ∈ R1×𝑛 hc, a residual transformation 𝐵 𝑙 ∈ R𝑛 hc ×𝑛 hc, and an output mapping𝐶 𝑙 ∈ R𝑛 hc ×1. The updateoftheresidualstateisthenformulatedas: 𝑋 𝑙+1 = 𝐵 𝑙 𝑋 𝑙 +𝐶 𝑙 F 𝑙 (𝐴 𝑙 𝑋 𝑙 ), (1) where F 𝑙 denotesthe 𝑙-thlayer(e.g.,anMoElayer),whoseinputandoutputshapesareboth R𝑑 . Notethattheactuallayerinput 𝐴 𝑙 𝑋 𝑙 ∈ R𝑑 isalso𝑑-dimensional,sotheexpandedresidual 7


widthdoesnotinfluencethedesignoftheinnerlayers. HCdecouplestheresidualwidthfrom the actual hidden size, offering a complementary scaling axis with minimal computational overhead,as𝑛 istypicallymuchsmallerthanthehiddensize𝑑. However,eventhoughHC hc has demonstrated potential in improving model performance, we find that the training will frequentlyexhibitnumericalinstabilitywhenstackingmultiplelayers,whichhindersthescaling ofHC. Manifold-ConstrainedResidualMapping. ThecoreinnovationofmHCistoconstrainthe residualmappingmatrix 𝐵 𝑙 tothemanifoldofdoublystochasticmatrices(theBirkhoffpolytope) M,andthusenhancethestabilityofsignalpropagationacrosslayers: 𝐵 𝑙 ∈ M ≔ {𝑀 ∈ R𝑛×𝑛 | 𝑀1𝑛 =1𝑛, 1 𝑇 𝑛 𝑀 =1 𝑇 𝑛 , 𝑀 ⩾ 0}. (2) Thisconstraintensuresthatthespectralnormofthemappingmatrix ∥𝐵 𝑙 ∥ 2 isboundedby1,so theresidualtransformationisnon-expansive,whichincreasesthenumericalstabilityduringboth theforwardpassandbackpropagation. Besides,thesetM isclosedundermultiplication,which guaranteesstabilityinthescenariosofdeepstacksofmHC.Inaddition,theinputtransformation 𝐴 𝑙 and output transformation 𝐶 𝑙 are also constrained to be non-negative and bounded via a Sigmoidfunctiontoavoidtheriskofsignalcancellation. DynamicParameterization. Theparametersofthreelinearmappingsaredynamicallygen- erated, which are decomposed into a dynamic (input-dependent) component and a static (input-independent)component. Giventheinput 𝑋 𝑙 ∈ R𝑛 hc ×𝑑 , itisfirstflattenedandnormal- ized: 𝑋ˆ 𝑙 =RMSNorm(vec(𝑋 𝑙 )) ∈ R1×𝑛 hc 𝑑 . Then,wefollowtheconventionalHCtogeneratethe unconstrainedrawparameters 𝐴˜ 𝑙 ∈ R1×𝑛 hc, 𝐵˜ 𝑙 ∈ R𝑛 hc ×𝑛 hc,and𝐶˜ 𝑙 ∈ R𝑛 hc ×1: 𝐴˜ 𝑙 =𝛼p 𝑙 re·(𝑋ˆ 𝑙 𝑊 𝑙 pre)+𝑆 𝑙 pre , (3) 𝐵˜ 𝑙 =𝛼r 𝑙 es·Mat(𝑋ˆ 𝑙 𝑊 𝑙 res)+𝑆 𝑙 res, (4) 𝐶˜ 𝑙 =𝛼p 𝑙 ost·(𝑋ˆ 𝑙 𝑊 𝑙 post)𝑇 +𝑆 𝑙 post , (5) where𝑊 𝑙 pre ,𝑊 𝑙 post ∈ R𝑛 hc 𝑑×𝑛 hc and𝑊 𝑙 res ∈ R𝑛 hc 𝑑×𝑛2 hc arelearnableparametersforgeneratingthe dynamic components; Mat(·) reshapes a vector of size 1×𝑛2 into a matrix of size 𝑛 ×𝑛 ; hc hc hc 𝑆pre ∈ R1×𝑛 hc,𝑆post ∈ R𝑛 hc ×1,and𝑆res ∈ R𝑛 hc ×𝑛 hc arelearnablestaticbiases;and𝛼pre ,𝛼res,𝛼post ∈ R 𝑙 𝑙 𝑙 𝑙 𝑙 𝑙 arelearnablegatingfactorsinitializedtosmallvalues. ApplyingParameterConstraints. Afterobtainingtheunconstrainedrawparameters 𝐴˜ 𝑙,𝐵˜ 𝑙,𝐶˜ 𝑙, wethenapplyconstraintsdescribedearliertothemtoenhancethenumericalstability. Tobe specific,fortheinputandoutputmappings,weemployaSigmoidfunction𝜎(·) toensuretheir non-negativityandboundedness: 𝐴 𝑙 =𝜎(𝐴˜ 𝑙 ), (6) 𝐶 𝑙 =2𝜎(𝐶˜ 𝑙 ). (7) Asfortheresidualmapping 𝐵˜ 𝑙,weprojectitontothemanifoldofdoublystochasticmatricesM. ThisisachievedbytheSinkhorn-Knoppalgorithm,whichfirstappliesanexponentialfunction to 𝐵˜ 𝑙 toensurepositivity,getting 𝑀(0) =exp(𝐵˜ 𝑙 ),andtheniterativelyperformscolumnandrow normalization: 𝑀(𝑡) =T 𝑟 (T 𝑐 (𝑀(𝑡−1))), (8) whereT 𝑟 andT 𝑐 denoterowandcolumnnormalization,respectively. Thisiterationconvergesto aconstraineddoublystochasticmatrix 𝐵 𝑙 =𝑀(𝑡 max ). Wechoose𝑡 max =20asapracticalvalue. 8


Shared Key-Value Multi-Query Attention Concatenation Lightning Indexer Sliding Window Selected KV Entries Compressed … KV Entries Index Scores Top-k Multi-Query … Selector Attention Compressed Compressed KV Entries … Indexer Keys … Indexer Queries Queries Token-Level Token-Level Compressor Compressor Hidden States of KV Tokens … Hidden State of Query Token Figure3 | CorearchitecturesofCSA.ItcompressesthenumberofKVentriesto 1 times,and 𝑚 thenapplies DeepSeekSparseAttention forfurtheracceleration. Additionally, asmallset of slidingwindowKVentriesiscombinedwiththeselectedcompressedKVentriestoenhance localfine-graineddependencies. 2.3. HybridAttentionwithCSAandHCA Asthecontextlengthreachesextremescales,theattentionmechanismemergesasthedominant computational bottleneck in a model. For DeepSeek-V4, we design two efficient attention architectures—CompressedSparseAttention(CSA)andHeavilyCompressedAttention(HCA) —andemploytheirinterleavedhybridconfiguration,whichsubstantiallyreducesthecompu- tationalcostofattentioninlong-textscenarios. CSAintegratesbothcompressionandsparse attention strategies: it first compresses the Key-Value (KV) cache of every 𝑚 tokens into one entry,andthenappliesDeepSeekSparseAttention(DSA)(DeepSeek-AI,2025)whereeachquery tokenattendstoonly𝑘compressedKVentries. HCAaimsforextremecompressionbyconsol- idatingtheKVcacheofevery𝑚′ (≫ 𝑚)tokensintoasingleentry. Thehybridarchitectureof CSAandHCAremarkablyimprovesthelong-contextefficiencyofDeepSeek-V4series,making one-million-token context feasible in practice. This subsection describes the core techniques ofourhybridattentionarchitecture,andwealsoprovideanopen-sourceimplementation1 to specifymoredetailsunambiguously. 2.3.1. CompressedSparseAttention ThecorearchitectureofCSAisillustratedinFigure3,whichfirstcompressestheKVcacheofeach 𝑚tokensintooneentry,andthenappliesDeepSeekSparseAttentionforfurtheracceleration. CompressedKey-ValueEntries. Let 𝐻 ∈ R𝑛×𝑑 beasequenceofinputhiddenstates,where 𝑛isthesequencelengthand𝑑 isthehiddensize. CSAfirstcomputestwoseriesofKVentries 𝐶𝑎 ,𝐶𝑏 ∈ R𝑛×𝑐 andtheircorrespondingcompressionweights 𝑍𝑎 ,𝑍𝑏 ∈ R𝑛×𝑐 ,where 𝑐 isthehead 1https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/tree/main/inference 9


dimension: 𝐶𝑎 = 𝐻·𝑊𝑎𝐾𝑉 , 𝐶𝑏 = 𝐻·𝑊𝑏𝐾𝑉 , (9) 𝑍𝑎 = 𝐻·𝑊𝑎𝑍 , 𝑍𝑏 = 𝐻·𝑊𝑏𝑍 , (10) where𝑊𝑎𝐾𝑉 ,𝑊𝑏𝐾𝑉 ,𝑊𝑎𝑍 ,𝑊𝑏𝑍 ∈ R𝑑×𝑐 aretrainableparameters. Next,each𝑚KVentriesin𝐶𝑎 and 𝐶𝑏 will be compressed into one entry according to their compression weights and learnable positionalbiases 𝐵𝑎 ,𝐵𝑏 ∈ R𝑚×𝑐 ,producing𝐶Comp ∈ R 𝑚 𝑛×𝑐 . Eachcompressedentry𝐶Comp ∈ R𝑐 is 𝑖 computedby [𝑆𝑎 ;𝑆𝑏 ] =Softmax ([𝑍𝑎 +𝐵𝑎 ;𝑍𝑏 +𝐵𝑏]), (11) 𝑚𝑖:𝑚(𝑖+1)−1 𝑚(𝑖−1):𝑚𝑖−1 row 𝑚𝑖:𝑚(𝑖+1)−1 𝑚(𝑖−1):𝑚𝑖−1 𝑚(𝑖+1)−1 𝑚𝑖−1 ∑︁ ∑︁ 𝐶Comp = 𝑆𝑎⊙𝐶𝑎+ 𝑆𝑏⊙𝐶𝑏 , (12) 𝑖 𝑗 𝑗 𝑗 𝑗 𝑗=𝑚𝑖 𝑗=𝑚(𝑖−1) where ⊙ denotestheHadamardproduct; Softmax (·) denotesthesoftmaxoperationalong row therowdimension,whichperformsnormalizationacrossthetotalof2𝑚elementsfromboth 𝑍𝑎 and 𝑍𝑏 . When 𝑖 =0, 𝑍𝑏 ispaddedwithnegativeinfinityand𝐶𝑏 ispadded 𝑚(𝑖−1):𝑚𝑖−1 𝑚(𝑖−1):𝑚𝑖−1 withzeros. Notethateach𝐶Comp isderivedfrom2𝑚KVentries,buttheindexesof𝐶𝑏 usedfor 𝑖 𝐶Comp andtheindexesof𝐶𝑎 usedfor𝐶Comp areoverlapped. Therefore,CSAinfactcompresses 𝑖 𝑖−1 thesequencelengthto 1 times. 𝑚 LightningIndexerforSparseSelection. AfterobtainingthecompressedKVentries𝐶Comp, CSAappliestheDSAstrategytoselecttop-kcompressedKVentriesforcoreattention. First, CSAperformsthesamecompressionoperationusedfor𝐶Comp togetcompressedindexerkeys 𝐾IComp ∈ R 𝑚 𝑛×𝑐𝐼 ,where𝑐𝐼 istheindexerheaddimension. Then,foraquerytoken𝑡,weproduce theindexerqueries {q 𝐼 ;q 𝐼 ;...;q 𝐼 } inalow-rankmanner: 𝑡,1 𝑡,2 𝑡,𝑛𝐼 ℎ c 𝑡 𝑄 =h𝑡 ·𝑊𝐷𝑄 , (13) [q 𝐼 ;q 𝐼 ;...;q 𝐼 ] =q 𝐼 =c 𝑄 ·𝑊𝐼𝑈𝑄 , (14) 𝑡,1 𝑡,2 𝑡,𝑛𝐼 𝑡 𝑡 ℎ where h𝑡 ∈ R𝑑 is the input hidden state of the query token 𝑡; c 𝑡 𝑄 ∈ R𝑑𝑐 is the compressed latentvectorforqueries;𝑑 𝑐 denotesthequerycompressiondimension;𝑛 ℎ 𝐼 denotesthenumber of indexer query heads; 𝑊𝐷𝑄 ∈ R𝑑×𝑑𝑐 and 𝑊𝐼𝑈𝑄 ∈ R𝑑𝑐×𝑐𝐼𝑛 ℎ 𝐼 are the down-projection and up- projectionmatricesforindexerqueries,respectively. Next,theindexscore 𝐼 𝑡,𝑠 ∈ R betweenthe querytoken𝑡 andaprecedingcompressedblock𝑠(𝑠<Floor( 𝑡 ))iscomputedby 𝑚 [𝑤 𝑡 𝐼 ,1 ;𝑤 𝑡 𝐼 ,2 ;...;𝑤 𝑡 𝐼 ,𝑛𝐼 ] =w 𝑡 𝐼 =h𝑡 ·𝑊𝑤 , (15) ℎ 𝑛𝐼 ∑︁ℎ (cid:16) (cid:17) 𝐼 𝑡,𝑠 = 𝑤 𝑡 𝐼 ,ℎ ·ReLU q 𝑡 𝐼 ,ℎ ·𝐾 𝑠 IComp , (16) ℎ=1 where𝑊𝑤 ∈ R𝑑×𝑛 ℎ 𝐼 isalearnablematrix;𝑤𝐼 ∈ R istheweightoftheℎ-thindexerhead. Fora 𝑡,ℎ querytoken𝑡,givenitsindexscores 𝐼 𝑡,: ,weemployatop-kselectortoselectivelyretainasubset ofcompressedKVentries CSprsComp forsubsequentcoreattention: 𝑡 (cid:110) (cid:12) (cid:111) C 𝑡 SprsComp = 𝐶 𝑠 Comp (cid:12) (cid:12) 𝐼 𝑡,𝑠 ∈ Top-k(𝐼 𝑡,: ) . (17) 10


Shared Key-Value Multi-Query Attention Concatenation Sliding Window Heavily KV Entries Compressed … KV Entries Queries Token-Level Compressor … Hidden States of KV Tokens Hidden State of Query Token Figure4 | CorearchitecturesofHCA.Itperformsheaviercompression,wheretheKVentriesof 𝑚′ (≫ 𝑚)tokenswillbeconsolidatedintoone. Also,weadditionallyintroduceasmallsetof slidingwindowKVentriestoenhancelocalfine-graineddependencies. Shared Key-Value MQA. After selecting the sparse KV entries, CSA then performs core attentioninaMulti-QueryAttention(MQA)(Shazeer,2019)manner,whereeachcompressed KVentryin CSprsComp servesasbothattentionkeyandvalue. Tobespecific,foraquerytoken𝑡, 𝑡 𝑄 wefirstproduceattentionqueries {q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛 ℎ } fromthecompressedlatentvectorc 𝑡 : [q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛 ℎ ] =q𝑡 =c 𝑡 𝑄 ·𝑊𝑈𝑄 , (18) where 𝑛 ℎ denotesthenumberofqueryheads;𝑊𝑈𝑄 ∈ R𝑑𝑐×𝑐𝑛 ℎ istheup-projectionmatricesfor 𝑄 queries. Notethatthelatentqueryvectorc issharedwiththatusedfortheindexerqueries. 𝑡 Next,weperformMQAon {q𝑡,𝑖 } and C 𝑡 SprsComp : (cid:16) (cid:17) o𝑡,𝑖 =CoreAttn query=q𝑡,𝑖,key=C 𝑡 SprsComp ,value=C 𝑡 SprsComp , (19) whereo𝑡,𝑖 ∈ R𝑐 isthecoreattentionoutputofthe𝑖-thheadatthe𝑡-thtoken;CoreAttn(·) denotes thecoreattentionoperation. GroupedOutputProjection. IntheconfigurationofDeepSeek-V4,𝑐𝑛 ℎisquitelarge. Therefore, directlyprojectingtheoutputsofthecoreattentionoperation [o𝑡,1 ;o𝑡,2 ;...;o𝑡,𝑛 ℎ ] =o𝑡 ∈ R𝑐𝑛 ℎ toa 𝑑-dimensionalhiddenstatewillimposeasubstantialcomputationalburden. Tomitigatethis cost, wedesignagroupedoutputprojectionstrategy. Tobespecific, wefirstsplit 𝑛 ℎ outputs into 𝑔 groups,andthenforeachgroupofoutputo 𝐺 𝑡,𝑖 ∈ R𝑐𝑛 𝑔 ℎ ,weprojectittoa𝑑 𝑔-dimensional intermediate output o 𝐺 𝑡,𝑖 ′ ∈ R𝑑𝑔, where 𝑑 𝑔 < 𝑐𝑛 𝑔 ℎ. Finally, we project the intermediate output [o 𝐺 𝑡,1 ′ ;o 𝐺 𝑡,2 ′ ;...;o 𝐺 𝑡,𝑔 ′] ∈ R𝑑𝑔𝑔 tothefinalattentionoutputoˆ𝑡 ∈ R𝑑 . 2.3.2. HeavilyCompressedAttention The core architecture of HCA is illustrated in Figure 4, which compresses the KV cache in a heaviermanner,butdoesnotemploysparseattention. CompressedKey-ValueEntries. Byandlarge,thecompressionstrategyofHCAissimilarto thatofCSA,butemploysalargercompressionrate𝑚′ (≫ 𝑚)anddoesnotperformoverlapped 11


compression. Let 𝐻 ∈ R𝑛×𝑑 be a sequence of input hidden states, HCA first computes the originalKVentries𝐶 ∈ R𝑛×𝑐 andtheircorrespondingcompressionweights𝑍 ∈ R𝑛×𝑐 : 𝐶 = 𝐻·𝑊𝐾𝑉 , (20) 𝑍 = 𝐻·𝑊𝑍 , (21) where𝑊𝐾𝑉 ,𝑊𝑍 ∈ R𝑑×𝑐 aretrainableparameters. Next,each𝑚′KVentriesin𝐶willbecompressed into one according to the compression weights and learnable positional biases 𝐵 ∈ R𝑚′×𝑐 , producing𝐶Comp ∈ R 𝑚 𝑛 ′×𝑐 . Eachcompressedentry𝐶Comp ∈ R𝑐 iscomputedby 𝑖 𝑆 𝑚′𝑖:𝑚′(𝑖+1)−1 =Softmax row (𝑍 𝑚′𝑖:𝑚′(𝑖+1)−1 +𝐵), (22) 𝑚′(𝑖+1)−1 ∑︁ 𝐶 𝑖 Comp = 𝑆 𝑗 ⊙𝐶 𝑗. (23) 𝑗=𝑚′𝑖 Throughthiscompressionoperation,HCAcompressesthesequencelengthto 1 times. 𝑚′ SharedKey-ValueMQAandGroupedOutputProjection. HCAalsoemploysthesharedKV MQAandgroupedoutputprojectionstrategiesasCSAdoes. AftertheKVcompression,fora querytoken𝑡,HCAfirstproducesattentionqueries {q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛 ℎ } inalow-rankmanner: c 𝑡 𝑄 =h𝑡 ·𝑊𝐷𝑄 , (24) [q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛 ℎ ] =q𝑡 =c 𝑡 𝑄 ·𝑊𝑈𝑄 , (25) whereh𝑡 ∈ R𝑑 istheinputhiddenstateofthequerytoken𝑡; 𝑛 ℎ denotesthenumberofquery heads;𝑊𝐷𝑄 ∈ R𝑑×𝑑𝑐 and𝑊𝑈𝑄 ∈ R𝑑𝑐×𝑐𝑛 ℎ arethedown-projectionandup-projectionmatricesfor queries,respectively. Next,weperformMQAon {q𝑡,𝑖 } and𝐶Comp: (cid:16) (cid:17) o𝑡,𝑖 =CoreAttn query=q𝑡,𝑖,key=𝐶Comp,value=𝐶Comp , (26) whereo𝑡,𝑖 ∈ R𝑐 isthecoreattentionoutputofthe𝑖-thheadatthe𝑡-thtoken. Next,asCSAdoes, HCAsplits𝑛 ℎ outputsinto𝑔 groups,andforeachgroupofoutputo 𝐺 𝑡,𝑖 ∈ R𝑐𝑛 𝑔 ℎ ,HCAprojectsit toa 𝑑 𝑔-dimensionalintermediateoutputo 𝐺 𝑡,𝑖 ′ ∈ R𝑑𝑔,where 𝑑 𝑔 < 𝑐𝑛 𝑔 ℎ. Finally,HCAprojectsthe intermediateoutput [o 𝐺 𝑡,1 ′ ;o 𝐺 𝑡,2 ′ ;...;o 𝐺 𝑡,𝑔 ′] ∈ R𝑑𝑔𝑔 tothefinalattentionoutputoˆ𝑡 ∈ R𝑑 . 2.3.3. OtherDetails InadditiontothecorearchitecturesofCSAandHCAdescribedabove, ourhybridattention incorporatesseveralothertechniques. Forwritingclarity,weomittheseadditionaltechniques fromtheaboveintroductionandwillbrieflydescribetheminthissubsection. Also,thissubsec- tionfocusesonlyonthecoreideasofthemandmayomitsometinydetailsforsimplicity. We encouragethereaderstorefertoouropen-sourceimplementationforunambiguousdetails. QueryandKey-ValueEntryNormalization. ForbothCSAandHCA,weperformanaddi- tionalRMSNormoperationoneachheadofthequeriesandtheonlyheadofthecompressedKV entries,justbeforethecoreattentionoperation. Thisnormalizationavoidsexplodingattention logitsandmayimprovetrainingstability. 12


PartialRotaryPositionalEmbedding. ForbothCSAandHCA,wepartiallyemploytheRotary PositionalEmbedding(RoPE)(Suetal.,2024)totheattentionqueries,KVentries,andthecore attentionoutputs. Tobespecific,foreachqueryvectorandKVentryvectorusedinCSAand HCA, we apply RoPE to its last 64 dimensions. Since the KV entries serve as both attention keysandvalues,thenaivecoreattentionoutputs {o𝑡,𝑖 } willcarryabsolutepositionembeddings, derivedfromtheweightedsumofKVentries. Asacountermeasure,wealsoapplyRoPEwith position−𝑖onthelast64dimensionsofeacho𝑡,𝑖. Inthisway,theoutputofthecoreattention willalsocarryrelativepositionembeddings—thecontributionofeachKVentrytothecore attentionoutputswillalsoberelatedtothedistancebetweenthequeryandtheKVentry. AdditionalBranchofSlidingWindowAttention. Inordertostrictlypreservecausalityin CSAandHCA,eachqueryattendstoonlyprecedingcompressedKVblocks. Consequently,a querycannotaccessinformationfromothertokenswithinitsowncompressedblock. Meanwhile, recenttokensusuallypossessgreaterrelevancetothequerytokeninlanguagemodeling. For thesereasons,weintroduceasupplementaryattentionbranchtobothCSAandHCAinasliding windowmanner,forbettermodelingoflocaldependencies. Tobespecific,foreachquerytoken, weadditionallyproduce𝑛 uncompressedKVentriescorrespondingtotherecent𝑛 tokens. win win In the core attention of CSA and HCA, these KV entries in the sliding window will be used alongwiththecompressedKVentries. Attention Sink. In the core attention of CSA and HCA, we employ the trick of attention sink (OpenAI, 2025; Xiao et al., 2024). To be specific, we set a series of learnable sink logits {𝑧′,𝑧′,...,𝑧′ }. For the ℎ-th attention head, Exp(𝑧′) will be added to the denominator of the 1 2 𝑛 ℎ ℎ attentionscore: 𝑠 ℎ,𝑖,𝑗 = (cid:205) 𝑘Exp E (𝑧 x ℎ p ,𝑖 ( ,𝑘 𝑧 ) ℎ + ,𝑖,𝑗 E ) xp(𝑧 ℎ ′) , (27) where 𝑠 ℎ,𝑖,𝑗,𝑧 ℎ,𝑖,𝑗 ∈ R denote the attention score and attention logit of the ℎ-th attention head betweenthe𝑖-thquerytokenandthe 𝑗-thprecedingtokenorcompressedblock. Thistechnique allowseachqueryheadtoadjustitstotalattentionscorestobenotequalto1,andeventobe near0. 2.3.4. EfficiencyDiscussion Due to the employment of hybrid CSA and HCA, together with low-precision computation andstorage,theattentionmoduleofDeepSeek-V4seriesachievesremarkableefficiencyinboth attention FLOPs and KV cache size, especially in long-context scenarios. First, we adopt a mixedstorageformatforKVentries: BF16precisionisusedfortherotarypositionalembedding (RoPE)dimensions,whileFP8precisionisappliedtotheremainingdimensions. Thishybrid representation reduces the KV cache size by nearly half compared with pure BF16 storage. Second, attention computation within the lightning indexer is performed in FP4 precision, which accelerates the attention operation under extremely long contexts. Third, relative to DeepSeek-V3.2,asmallerattentiontop-kischoseninDeepSeek-V4series,therebyimproving modelefficiencyonshort-andmedium-lengthtexts. Finally,andmostimportantly,compressed attentionandhybridattentiontechniquessubstantiallyreduceboththeKVcachesizeandthe computationalFLOPs. TakingBF16GQA8(Ainslieetal.,2023)withaheaddimensionof128asthebaseline—one ofthecommonconfigurationsofLLMattention—theKVcachesizeofDeepSeek-V4seriescan bedramaticallyreducedtoapproximately2%timesofthatbaselineinthe1M-contextsetting. 13


Algorithm1MuonOptimizerforDeepSeek-V4 Require: Learningrate𝜂,momentum 𝜇,weightdecay 𝜆,updaterescalingfactor𝛾 1: foreachtrainingstep𝑡 do 2: foreachlogicallyindependentweight𝑊 ∈ R𝑛×𝑚 do 3: 𝐺 𝑡 =∇ 𝑊 L 𝑡 (𝑊 𝑡−1 ) ⊲Computegradients 4: 𝑀 𝑡 =𝜇𝑀 𝑡−1 +𝐺 𝑡 ⊲Accumulatemomentumbuffer 5: 𝑂 𝑡 ′ =HybridNewtonSchulz(𝜇𝑀 𝑡 +𝐺 𝑡 ) ⊲NesterovtrickandhybridNewton-Schulz √︁ 6: 𝑂 𝑡 =𝑂 𝑡 ′· max(𝑛,𝑚)·𝛾 ⊲RescaletheupdateRMS 7: 𝑊 𝑡 =𝑊 𝑡−1 ·(1−𝜂𝜆)−𝜂𝑂 𝑡 ⊲Performweightdecayandupdate 8: endfor 9: endfor Moreover,evenwhencomparedwithDeepSeek-V3.2(DeepSeek-AI,2025)—alreadyanefficient baseline—DeepSeek-V4seriesstillexhibitssubstantialadvantagesinefficiency. Acomparison oftheirinferenceFLOPsandKVcachesizeisprovidedintherightpartofFigure1. 2.4. MuonOptimizer WeemploytheMuon(Jordanetal.,2024;Liuetal.,2025)optimizerforthemajorityofmodules inDeepSeek-V4seriesduetoitsfasterconvergenceandimprovedtrainingstability. Thefull algorithmofourMuonoptimizationissummarizedinAlgorithm1. BasicConfigurations. WemaintaintheAdamW(LoshchilovandHutter,2017)optimizerfor the embedding module, the prediction head module, the static biases and gating factors of mHCmodules,andtheweightsofallRMSNormmodules. Allothermodulesareupdatedwith Muon. FollowingLiuetal.(2025), wealsoapplyweightdecaytoMuonparameters, usethe Nesterov(Jordanetal.,2024;Nesterov,1983)trick,andrescaletheRootMeanSquare(RMS)of theupdatematrixforreutilizationofourAdamWhyper-parameters. Differentfromthem,we usehybridNewton-Schulziterationsfororthogonalization. HybridNewton-SchulzIterations. Foragivenmatrix 𝑀,letitsSingularValueDecomposition (SVD)be 𝑀 =𝑈Σ𝑉𝑇 . TheNewton-Schulziterationsaimtoapproximatelyorthogonalize 𝑀 tobe 𝑈𝑉𝑇 . Usually, 𝑀 willbefirstnormalizedas 𝑀 0 = 𝑀/||𝑀|| 𝐹 toensureitsmaximumsingularvalue doesnotexceed1. Then,eachNewton-Schulziterationperformsthefollowingoperation: 𝑀 𝑘 =𝑎𝑀 𝑘−1 +𝑏(𝑀 𝑘−1 𝑀 𝑘 𝑇 −1 )𝑀 𝑘−1 +𝑐(𝑀 𝑘−1 𝑀 𝑘 𝑇 −1 )2𝑀 𝑘−1 . (28) OurhybridNewton-Schulzperforms10iterationsovertwodistinctstages. Duringthefirst8 steps,weusecoefficients (𝑎,𝑏,𝑐) = (3.4445,−4.7750,2.0315) todriverapidconvergence,bringing thesingularvaluescloseto1. Inthefinal2steps,weswitchtocoefficients (𝑎,𝑏,𝑐) = (2,−1.5,0.5), whichstabilizethesingularvaluespreciselyat1. AvoidingExplodingAttentionLogits. TheattentionarchitectureofDeepSeek-V4seriesal- lowsustodirectlyapplyRMSNormontheattentionqueriesandKVentries,whicheffectively preventsattentionlogitsfromexploding. Consequently,wedonotemploytheQK-Cliptech- nique(Liuetal.,2025)inourMuonoptimizer. 14


  1. General Infrastructures 3.1. Fine-GrainedCommunication-ComputationOverlapinExpertParallelism Mixture-of-Experts (MoE) can be accelerated via Expert Parallelism (EP). However, EP re- quirescomplexinter-nodecommunicationandimposessubstantialdemandsoninterconnect bandwidthandlatency. ToalleviatethecommunicationbottleneckinEPandachievehigher end-to-end performance under lower interconnection bandwidth requirements, we propose afine-grainedEPschemethatfusescommunicationandcomputationintoasinglepipelined kernelforcommunication-computationoverlapping. Communication Latency Can Be Hidden. The key insight of our EP scheme is that the communicationlatencycanbeeffectivelyhiddenbeneathcomputationinMoElayers. Asshown inFigure5,inDeepSeek-V4series,eachMoElayercanbedecomposedmainlyintofourstages: twocommunication-boundstages,DispatchandCombine,andtwocomputation-boundstages, Linear-1 and Linear-2. Our profiling reveals that within a single MoE layer, the total time of communicationislessthanthatofthecomputation. Therefore,afterfusingcommunicationand computationintoaunifiedpipeline,computationremainsthedominantbottleneck,implying that the system can tolerate lower interconnect bandwidth without degrading end-to-end performance. (a) Naive Solution L1 Act L2 (b) Comet Communication Theoretical speedup: 1.42× Computation L1 Act L2 (c) Ours Dispatch Dispatch All-to-All Computation L1 L2 L1 L2 L1 L2 Theoretical speedup: 1.92× Linear 1 GEMM SwiGLU + FP8 Cast Activation Act Act Act Linear 2 GEMM & Combine Combine All-to-All Expert Wave 1 Expert Wave 2 Expert Wave 3 Figure5|IllustrationofourEPschemewithrelatedworks. Comet(Zhangetal.,2025b)overlaps DispatchwithLinear-1,andLinear-2withCombine,separately. OurEPschemeachievesafiner- grainedoverlappingbysplittingandschedulingexpertsintowaves. Thetheoreticalspeedupis evaluatedintheconfigurationoftheDeepSeek-V4-Flasharchitecture. Fine-Grained EP Scheme. To further lower the interconnect bandwidth requirement and amplifythebenefitsofoverlapping,weintroduceafiner-grainedexpertpartitioningscheme. Inspiredbymanyrelatedworks(Aimuyoetal.,2025;Zhangetal.,2025b),wesplitandschedule theexpertsintowaves. Eachwaveconsistsofasmallportionofexperts. Assoonasallexperts withinthewavehavecompletedtheircommunication,computationcancommenceimmediately withoutwaitingforotherexperts. Insteadystate,computationofcurrentwave,tokentransferfor thenextwave,andresultsendingofcompletedexpertsallproceedconcurrently,asdemonstrated inFigure5. Thisformsafine-grainedpipelineamongexperts,keepingbothcomputationand communicationcontinuousthroughoutthewave. Thewave-basedschedulingspeedsupthe 15

performance on extreme cases such as Reinforcement Learning (RL) rollout, which usually encounterslong-tailsmallbatches. Performance and Open-Sourced Mega-Kernel. We validated the fine-grained EP scheme on both NVIDIA GPUs and HUAWEI Ascend NPUs platforms. Compared against strong non-fused baselines, it achieves 1.50 ∼ 1.73× speedup for general inference workloads, and upto1.96×forlatency-sensitivescenariossuchasRLrolloutsandhigh-speedagentserving. Wehaveopen-sourcedtheCUDA-basedmega-kernelimplementationnamedMegaMoE2 asa componentofDeepGEMM. ObservationsandProposals. Weshareobservationsandlessonsfromkerneldevelopment andoffersomeproposalstohardwarevendors,inthehopeofaidingefficienthardwaredesign andachievingbettersoftware-hardwareco-design: • Computation-CommunicationRatio. Fullcommunication-computationoverlaphinges onthecomputation-communicationratio,ratherthanthebandwidthsolely. Denotingpeak computethroughputas𝐶 andinterconnectbandwidthas 𝐵,communicationcanbefully hiddenwhen𝐶/𝐵 ⩽ 𝑉 /𝑉 ,where𝑉 denotesthecomputationvolumeand𝑉 comp comm comp comm refers to the communication volume. For DeepSeek-V4-Pro, where each token-expert pairrequires6ℎ𝑑 FLOPs(SwiGLUgate,up,anddownprojections)butonly3ℎbytesof communication(FP8Dispatch+BF16Combine),thissimplifiesto: 𝐶 ⩽ 2𝑑 =6144 FLOPs/Byte. 𝐵 Thatis,eachGBpsofinterconnectbandwidthsufficestohidethecommunicationfor6.1 TFLOP/sofcompute. Oncebandwidthmeetsthisthreshold,itceasestobethebottleneck, and devoting additional silicon area to further bandwidth brings diminishing returns. We encourage future hardware designs to target such balance points rather than scale bandwidthunconditionally. • Power Budget. Extreme kernel fusion drives compute, memory, and network to high loadsimultaneously,makingpowerthrottlingakeyperformancelimiter. Wesuggestthat future hardware designs provide sufficient power headroom for such fully concurrent workloads. • CommunicationPrimitives. Weadoptapull-basedapproachwhereeachGPUactively reads data from remote GPUs, avoiding the high notification latency that fine-grained pushentails. Futurehardwarewithlower-latencycross-GPUsignalingwouldmakepush viableandenablemorenaturalcommunicationpatterns. • ActivationFunction. WeproposereplacingSwiGLUwithalow-costelement-wiseactiva- tionthatinvolvesnoexponentialordivisionoperations. Thislightensthepost-GEMM processingdirectly,andunderthesameparameterbudget,removingthegateprojection enlargestheintermediatedimension𝑑,furtherrelaxingthebandwidthrequirement. 3.2. FlexibleandEfficientKernelDevelopmentwithTileLang Inpractice,ourelaboratemodelarchitecturewouldhaveresultedinhundredsoffine-grained TorchATenoperators. WeadoptTileLang(Wangetal.,2026)todevelopasetoffusedkernels to replace the vast majority of them, delivering optimal performance with minimal effort. It 2https://github.com/deepseek-ai/DeepGEMM/pull/304 16


alsoallowsustoquicklyprototypeoperatorslikeattentionvariantsduringvalidation. These kernelsplaycriticalrolesinmodelarchitecturedevelopment,large-scaletraining,andultimately productiondeploymentofinferenceservices. AsaDomain-SpecificLanguage(DSL),TileLang balancesdevelopmentproductivitywithruntimeefficiency,enablingrapiddevelopmentwhile supporting deep, iterative optimizations within the same codebase. Additionally, we collab- oratecloselywiththeTileLangcommunitytofosteramoreagile, efficient, andstablekernel developmentworkflow. Reducing Invocation Overhead with Host Codegen. As accelerators continue to grow in performance, CPU-side orchestration overhead becomes increasingly prominent. For small, highlyoptimizedkernels,suchfixedhostoverheadcaneasilycaputilizationandthroughput. Acommonsourceofthisoverheadisthathost-sidelogic,suchasruntimecontractchecks,is typicallywritteninPythonforflexibilityandthusincursafixedper-invocationcost. WemitigatethisoverheadwithHostCodegen,whichmovesmosthost-sidelogicintogen- erated host code. Specifically, we first co-generate the device kernel and a lightweight host launcherattheIR(IntermediateRepresentation)level,embeddingthenecessarymetadata—such asdatatypes,rank/shapeconstraints,andstride/layoutassumptions—parsedfromthelan- guage frontend. The launcher is then lowered to the host source code built on top of the TVM-FFI(Chenetal.,2018)framework,whosecompactcallingconventionandzero-copytensor interoptogetherminimizehost-sideoverhead. Atruntime,thisgeneratedhostcodeperforms validationandargumentmarshaling,shiftingallper-invocationchecksoutofthePythonexe- cutionpath. OurmeasurementsshowthatCPU-sidevalidationoverheaddropsfromtensor hundredsofmicrosecondstolessthanonemicrosecondperinvocation. SMT-Solver-Assisted Formal Integer Analysis. TileLang kernels involve complex tensor indexarithmeticthatrequiresstrongformalintegeranalysis. Duringcompilationpassessuch aslayoutinference,memoryhazarddetection,andboundanalysis,thecompilermustverify whetherintegerexpressionssatisfyspecificpropertiestoenablethecorrespondingoptimiza- tions. Therefore,strongerformalanalysiscapabilitiescanunlockmoreadvancedandcomplex optimizationopportunities. Tothisend,weintegratetheZ3SMTsolver(DeMouraandBjørner,2008)intoTileLang’s algebraicsystem,providingformalanalysiscapabilityformostintegerexpressionsintensor programs. Westrikeabalancebetweencomputationaloverheadandformalexpressivenessby translatingTileLang’sintegerexpressionsintoZ3’squantifier-freenon-linearintegerarithmetic (QF_NIA).BasedonIntegerLinearProgramming(ILP)solvers,QF_NIAseamlesslyresolves standardlinearintegerexpressionscommoninkernels. Furthermore,itsinherentnon-linear reasoningcapacityeffectivelyaddressesadvancedchallengeslikevectorizationovervariable tensorshapes. Underreasonableresourcelimits,Z3elevatesoveralloptimizationperformance while restricting compilation time overhead to just a few seconds. The impact is substantial acrossmultiplepasses,includingvectorization,barrierinsertion,andcodesimplification. NumericalPrecisionandBitwiseReproducibility. Inproductionsettings,numericalcorrect- nessandreproducibilityareascriticalasrawthroughput. Wethereforeprioritizeaccuracyby default: fast-mathoptimizationsaredisabledatthecompilerlevel,andprecision-affectingap- proximationsareprovidedonlyasexplicit,opt-infrontendoperators(e.g.,T.__exp,T.__log, and T.__sin). Conversely, when strict IEEE-754 semantics are required, TileLang provides 17


IEEE-compliantintrinsicswithexplicitroundingmodes(e.g.,T.ieee_fsqrt,T.ieee_fdiv, andT.ieee_add),enablingdeveloperstopreciselyspecifynumericalbehavior. Wealsotargetbitwisereproducibilityforvalidatingkernelsagainsthand-writtenCUDA baselines. We align TileLang’s algebraic simplification and lowering rules with mainstream CUDAtoolchains(e.g.,NVCC)toavoidtransformationsthatintroduceunintendedbit-level differences. Layoutannotations(e.g.,T.annotate_layout)furtherallowuserstopindown layout-dependentloweringdecisions,keepingevaluationandaccumulationorderconsistent withthereferenceCUDAimplementationandthusenablingbit-identicaloutputswhendesired. Ourevaluationshowsthattheseaccuracy-andreproducibility-orienteddesignchoicesdo notsacrificeperformance: underconservativedefaults,TileLangkernelsremaincompetitive, whileexposingknobstoselectivelyrelaxnumericalconstraintsforhigherspeed. 3.3. High-PerformanceBatch-InvariantandDeterministicKernelLibraries Toenableefficienttrainingandinference,wedevelopacomprehensivesetofhigh-performance computational kernels. Beyond basic functionalities and maximizing hardware utilization, anotherpivotaldesigngoalistoensuretrainingreproducibilityandbitwisealignmentamong pre-training, post-training, and inference pipelines. Therefore, we implement end-to-end, bitwisebatch-invariant,anddeterministickernelswithminimalperformanceoverhead. These kernelsarehelpfulfordebugging,stabilityanalysis,andconsistentpost-trainingbehavior. BatchInvariance. Batchinvarianceensuresthattheoutputofanygiventokenremainsbitwise identical,regardlessofitspositionwithinabatch. Toimplementbatchinvariance,theprimary challengesarelistedasfollows: • Attention. Toachievebatchinvariance, wecannotusethesplit-KVmethod(Daoetal., 2023),whichdistributestheattentioncomputationforasinglesequenceacrossmultiple Stream Multiprocessors (SMs) to balance the load of SMs. However, abandoning this techniquewillleadtoseverewave-quantizationproblems3,whichcanadverselyaffect GPUutilization. Toaddressthis,wedevelopadual-kernelstrategyforbatch-invariant decoding. Thefirstkernelcomputestheattentionoutputforanentiresequencewithin asingleSM,ensuringhighthroughputforfullyoccupiedwaves. Thesecondkernel,to minimizethelatencyofthefinalpartially-filledwaveandthusalleviatewave-quantization, uses multiple SMs for a single sequence. For the bitwise identity of these two kernels, wecarefullydesignthecalculationpathofthesecondkerneltoensureitsaccumulation orderisthesameasthatofthefirstkernel. Additionally,thesecondkernelutilizesdis- tributedsharedmemory4 withinthread-blockclusters,enablinghigh-speeddataexchange acrossSMs. Thisdual-kernelmethodeffectivelyconfinestheoverheadofbatch-invariant decodingtobenegligible. • MatrixMultiplication. TraditionalcuBLASlibrary(NVIDIACorporation,2024)cannot achieve batch invariance. Therefore, we replace it end-to-end with DeepGEMM (Zhao etal.,2025). Furthermore,forverysmallbatchsizes,conventionalimplementationusually employssplit-k(Osamaetal.,2023)techniquestoimproveperformance. Unfortunately, split-ktechniquescannotguaranteebatchinvariance,apivotalfeatureinDeepSeek-V4. 3https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplicat ion/index.html#wave-quant 4https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/writing-cuda-kernels .html#distributed-shared-memory 18


Therefore,weabandonsplit-kinmostscenarios,which,however,maycauseperformance degradation. Toaddressthis,weintroduceasetofoptimizationsthatenableourimple- mentationofmatrixmultiplicationtomatchorevensurpasstheperformanceofstandard split-kinmostmajorscenarios. Determinism. Deterministictrainingishighlybeneficialfordebugginghardwareorsoftware issues. Moreover,whentrainingexhibitsanomaliessuchaslossspikes,determinismenables researcherstomoreeasilypinpointnumericalcausesandfurtherrefinethemodeldesign. Non- determinismintrainingtypicallystemsfromnon-deterministicaccumulationorder,oftendue totheuseofatomicadditioninstructions. Thisissueprimarilyoccursduringthebackwardpass, notablyatthefollowingparts: • Attention Backward. In conventional implementations of backward propagation for sparse attention, we use atomicAdd to accumulate gradients for the KV tokens. This introducesnon-determinismduetothenon-associativityoffloating-pointaddition. To addressthisproblem,weallocateseparateaccumulationbuffersforeachSM,followedby aglobaldeterministicsummationacrossallbuffers. • MoE Backward. When multiple SMs from different ranks concurrently write data to the same buffer on a receiving rank, negotiating writing positions also introduces non- determinism. Toresolvethis,wedesignatokenorderpre-processingmechanismwithin each single rank, combined with buffer isolation across multiple ranks. This strategy ensuresdeterminismofboththesendresultsofexpertparallelismandtheaccumulation orderintheMoEbackwardpass. • MatrixMultiplicationinmHC.mHCinvolvesamatrixmultiplicationwithanoutputdi- mensionofonly24. Forverysmallbatchsizes,wearecompelledtousethesplit-k(Osama et al., 2023) algorithm, whose naive implementation will cause non-determinism. To overcomethis,weoutputeachsplitpartseparatelyandperformadeterministicreduction inasubsequentkernel,therebypreservingbothperformanceanddeterminism. 3.4. FP4Quantization-AwareTraining Toachieveinferenceaccelerationandmemorysavingsatdeployment,weintroduceQuantization- Aware Training (QAT) (Jacob et al., 2018) during the post-training stage, enabling the model to adapt to the precision degradation introduced by quantization. We apply FP4 (MXFP4) quantization (Rouhani et al., 2023) to two components: (1) MoE expert weights, which are a major source of GPU memory occupancy (OpenAI, 2025), and (2) the Query-Key (QK) path in the indexer of CSA, where QK activations are cached, loaded, and multiplied entirely in FP4,acceleratingattentionscorecomputationinlong-contextscenarios. Inaddition,wefurther quantize the index scores 𝐼 from FP32 to BF16 during this QAT process. This optimization :,: achievesa2×speedupforthetop-kselector,whilepreservinga99.7%recallrateofKVentries. ForMoEexpertweights,followingthecommonpracticeofQAT,theFP32masterweights maintained by the optimizer are first quantized to FP4, then dequantized back to FP8 for computation. Notably,ourFP4-to-FP8dequantizationislossless. ThisisbecauseFP8(E4M3) has 2 additional exponent bits compared with FP4 (E2M1), offering a larger dynamic range. Consequently,aslongastheratiobetweenthemaximumandminimumscalefactorsoftheFP4 sub-blocks (1×32 tiles) within each FP8 quantization block (128×128 tiles) does not exceed acertainthreshold,thefine-grainedscaleinformationcanbefullyabsorbedbytheextended dynamicrangeofFP8. Weempiricallyverifythatcurrentweightssatisfythiscondition. This allows the entire QAT pipeline to fully reuse the existing FP8 training framework without 19


anymodification. Inthebackwardpass,gradientsarecomputedwithrespecttothesameFP8 weightsintheforwardpassanddirectlypropagatedbacktotheFP32masterweights,equivalent toapplyingtheStraight-ThroughEstimator(STE)throughthequantizationoperation. Thisalso avoidstheneedtore-quantizetransposedweights. During the inference and rollout phases of RL training, which do not involve backward passes, we directly use real FP4 quantized weights instead of simulated quantization. This ensuresthatmodelbehaviorduringsamplingisfullyconsistentwithonlinedeployment,while alsoreducingkernelmemoryloadingforactualspeedupandsignificantlyloweringmemory consumption. WeprocesstheQKpathintheindexerofCSAsimilarly. 3.5. TrainingFramework Our training framework is built upon the scalable and efficient infrastructure developed for DeepSeek-V3(DeepSeek-AI,2024). IntrainingDeepSeek-V4,weinheritthisrobustfoundation whileintroducingseveralkeyinnovationstoaccommodateitsnovelarchitecturalcomponents— specificallytheMuonoptimizer,mHC,andthehybridattentionmechanism—whilemaintaining hightrainingefficiencyandstability. 3.5.1. EfficientImplementationofMuon TheMuonoptimizerrequiresthefullgradientmatrixtocomputeparameterupdates,which presentsachallengewhencombinedwiththeZeroRedundancyOptimizer(ZeRO)(Rajbhandari etal.,2020). TraditionalZeROisdesignedforelement-wiseoptimizerslikeAdamW,wherea singleparametermatrixcanbepartitionedandupdatedacrossmultipleranks. Toaddressthis conflict,wedesignahybridstrategyofZeRObucketassignmentforMuon. For dense parameters, we limit the maximum size of ZeRO parallelism and employ a knapsackalgorithmtoassignparametermatricestotheseranks,ensuringeachrankmanagesa roughlybalancedload. Thebucketoneachrankispaddedtomatchthesizeofthelargestbucket acrossranks,facilitatingefficientreduce-scatteroperations. Thispaddingtypicallyincursless than10%memoryoverheadinoursetup,whereeachrankmanagesnomorethanfiveparameter matrices. When the overall size of data parallelism exceeds the limit for ZeRO, we compute theMuonupdateredundantlyacrosstheextradata-parallelgroups,tradingcomputationfor reducedtotalbucketmemory. For MoE parameters, we optimize each expert independently. We first flatten all down projection matrices in SwiGLU (Shazeer, 2020) of all experts across all layers, followed by flattenedupprojectionmatricesandgatematrices. Then,wepadtheflattenedvectortoensure wecanevenlydistributethisvectoracrossallrankswithoutsplittinganylogicallyindependent matrix. Giventhelargenumberofexperts,wedonotimposealimitofZeROparallelismfor MoEparameters,andthepaddingoverheadisalsonegligible. Additionally,oneachrank,consecutiveparametersofidenticalshapewillbeautomatically merged,enablingbatchedexecutionoftheNewton-Schulziterationsforbetterhardwareutiliza- tion. Furthermore,weobservethattheNewton-SchulziterationsinMuonremainstablewhen computedwithBF16matrixmultiplications. Leveragingthis,wefurtherquantize,inastochastic roundingmanner,theMoEgradientstobesynchronizedacrossdata-parallelrankstotheBF16 precision, halving the communication volume. To avoid accumulation errors introduced by low-precisionadders,wereplaceconventionaltree-orring-basedreduce-scattercollectiveswith atwo-phaseapproach. First,anall-to-alloperationexchangeslocalgradientsacrossranks,and theneachrankperformsalocalsuminFP32. Thisdesignmaintainsnumericalrobustness. 20


3.5.2. Cost-EffectiveandMemory-EfficientImplementationofmHC TheintroductionofmHCincreasesbothactivationmemoryconsumptionandcommunication volumebetweenpipelinestages,comparedwithconventionalresidualconnections. Tomitigate thesecosts,weimplementseveraloptimizationstrategies. Firstly, we carefully design and implement fused kernels of mHC for both training and inference. Secondly,weintroducearecomputationstrategythatselectivelycheckpointsinterme- diatetensors. Specifically,werecomputemosthiddenstatesbetweenlayersandallnormalized layerinputs,whileavoidingrecomputationofcompute-intensiveoperations. Thisachievesa balancebetweenmemorysavingandcomputationaloverhead. Thirdly,weadjusttheDualPipe 1F1Boverlappingschemetoaccommodatetheincreasedpipelinecommunicationandenable concurrentexecutionofsomeoperationsinmHC. Collectively,theseoptimizationsconstrainthewall-timeoverheadofmHCtoonly6.7%of theoverlapped1F1Bpipelinestage. Moredetailsoftheengineeringoptimizationcanbefound inthededicatedmHCpaper(Xieetal.,2026). 3.5.3. ContextualParallelismforLong-ContextAttention Conventional Context Parallelism (CP) partitions the sequence dimension, with each rank maintainingcontiguous𝑠tokens. Thisintroducestwochallengestoourcompressedattention mechanisms(i.e.,CSAandHCA).Ontheonehand,trainingsamplesarepackedfrommultiple sequences,andeachsequenceiscompressedindependentlybyafactorof𝑚(or𝑚′),withany trailingtokensfewerthan 𝑚 beingdiscarded. Consequently, thecompressedKVlengthsare typically less than 𝑠 and vary across ranks. On the other hand, the compression requires 𝑚 𝑚 consecutiveKVentries,whichmaystraddletheboundarybetweentwoneighboringCPranks. Toaddressthesechallenges,wedesignatwo-stagecommunicationapproach. Inthefirst stage, each rank 𝑖 sends its last 𝑚 uncompressed KV entries to rank 𝑖+1. Then, rank 𝑖+1 compressessomeofthesereceivedentriestogetherwithitslocal 𝑠uncompressedKVentries, producingafixedlengthof 𝑠 +1compressedentries,inwhichexistsomepaddingentries. In 𝑚 thesecondstage,anall-gatheroperationacrossallCPrankscollectsthelocallycompressedKV entries. Then,afusedselect-and-padoperatorreorganizesthemintothefullsetofcompressed KVentrieswithatotallengthofcp_size· 𝑠 . Anypaddingentriesareplacedatthetail. For 𝑚 HCAandtheindexerinCSA,thevisiblerangeofcompressedKVentriesforeachquerytoken can be precomputed by rules. For the sparse attention in CSA, the top-𝑘 selector explicitly specifiestheindicesofvisiblecompressedKVentriesforeachquery. 3.5.4. ExtendedAutomaticDifferentiationforFlexibleActivationCheckpointing Conventionalactivationcheckpointingimplementationsoperateatthegranularityofanentire module,decidingwhethertoretainorrecomputeitsoutputactivationsduringthebackward pass. Thiscoarsegranularityoftenleadstosuboptimaltrade-offsbetweenrecomputationcost andactivationmemoryfootprint. Analternativeapproachistomanuallyimplementtheforward andbackwardlogicofanentirelayer,explicitlymanagingtensorcheckpointingstates. While enablingfine-grainedcontrol,thismethodlosestheconvenienceoftheautomaticdifferentiation framework,substantiallyincreasingdevelopmentcomplexity. Toachievefine-grainedcontrolwithoutsacrificingprogrammingefficiency,weimplementa tensor-levelactivationcheckpointingmechanismwithautomaticdifferentiationsupport. With thismechanism,developersonlyneedtoimplementtheforwardpassandselectivelyannotate 21


individualtensorsforautomaticcheckpointingandrecomputation. Ourframeworkleverages TorchFX(Reedetal.,2022)totracethefullcomputationgraph. Foreachannotatedtensor, it performsabackwardtraversaltoidentifytheminimalsubgraphrequiredforitsrecomputation. Wedefinetheseminimalsubgraphsasrecomputationgraphsandinsertthemintothebackward logicjustbeforethecorrespondinggradientcomputation. Comparedwiththemanualimplementation,thisdesignintroducesnoadditionaloverhead during training. Recomputation in this framework is implemented by directly freeing the GPU memory of the annotated tensor and reusing the storage pointer from the recomputed tensor,withoutanyGPUmemorycopy. Furthermore,sincegraphtracingexecutesthemodel concretely,wecantracktheunderlyingstoragepointerofeachtensor,whichenablesautomatic deduplicationofrecomputationfortensorsthatsharestorage(e.g.,theinputandoutputofa reshapeoperation). Thisrelievesdevelopersfromreasoningaboutlow-levelmemorydetails whenannotatingrecomputation. 3.6. InferenceFramework OurinferenceframeworklargelyinheritsfromthatofDeepSeek-V3,withsomedifferencesin KVCachemanagement. 3.6.1. KVCacheStructureandManagement ToefficientlymanagetheheterogeneousKVcachesarisingfromthehybridattentionmechanism inDeepSeek-V4,wedesignacustomizedKVcachelayout. ThelayoutisillustratedinFigure6, andwewillelaborateonitindetailasfollows. HeterogeneousKVEntriesinDeepSeek-V4. ThehybridattentionmechanisminDeepSeek- V4 series introduces multiple types of KV entries with different Key-Value (KV) cache sizes andupdaterules. Thelightningindexerforsparseselectionintroducesadditionaldimensions into the KV cache that possess embedding sizes distinct from those in the primary attention. ThecompressiontechniquesemployedinCSAandHCAreducethesequencelengthbyfactors of 1 and 1 ,respectively,therebydecreasingtheoverallKVcachesize. Asaresult,KVcache 𝑚 𝑚′ sizesvaryacrossdifferentlayers. Furthermore,SlidingWindowAttention(SWA)layersalso operate with distinct KV cache sizes, as well as separate cache hit and eviction policies. In the compression branch, one KV entry is generated for every 𝑚 tokens. When the number of remaining tokens is insufficient for compression, all pending tokens and their associated hidden states must be retained in a buffer until the compression operation can be executed. Thesebufferedtokensrepresentasequencestatedeterminedbypositionalcontextandarealso managedwithintheKVcacheframework. Challenges in Managing Hybrid Attention KV Cache. The hybrid attention mechanism violatesfundamentalassumptionsbehindPagedAttentionanditsvariants. Althoughrecent hybridKVcachemanagingalgorithms(e.g.,Jenga(Zhangetal.,2025a),Hymba(Dongetal., 2025)) target general hybrid attention models or specific structures, two principal obstacles preventconsolidatingKVcachesacrossalllayersunderthePagedAttentionframework: • Diversecachepolicies,suchasthoseusedinSlidingWindowAttention. • Constraintsimposedbyhigh-performanceattentionkernels,includingalignmentrequire- ments. 22


State Cache KV Cache Uncompressed Request 1 SWA KV Block 0 KV State Uncompressed Request 2 SWA KV Block 1 KV State Uncompressed Request 3 SWAL aKyVer-0 SWA KV Layer-2 CSA State Block 2 KV State …UncompresseLdayer-3 HCA State SWA KV KV State Layer-n SWA KV … Uncompressed Request R SWA KV Block N KV State …… CSA KV HCA KV LaCySeAr- 2KV CSA Indexer KV CSA Main KV HCA KV of k1 tokens of k1 tokens La C y S e A r- 3 KV HCA KV of k2 tokens LaHyCeAr- 4KV CSA Indexer KV CSA Main KV CSA KV of k1 tokens of k1 tokens La H y C e A r- 5 KV HCA KV of k2 tokens ... ... CSA KV HCA KV …… Figure6 | IllustrationoftheKVcacheLayoutforDeepSeek-V4. TheKVcacheisorganizedinto twoprimarycomponents: aclassicalKVcacheforCSA/HCA,andastatecacheforSWAand unready-for-compressiontokensinCSA/HCA.Inthestatecache,eachrequestisassigneda fixed-sizecacheblock. Withinthisblock,theSWAsegmentstorestheKVentriescorresponding tothemostrecent𝑛 tokens,whiletheCSA/HCAsegmentstoresuncompressedtailstates win thatarenotyetreadyforcompression. IntheclassicalKVcache,weallocatemultipleblocks perrequest. Eachcacheblockcoverslcm(𝑚,𝑚′) originaltokens,producing𝑘 = lcm(𝑚,𝑚′) CSA 1 𝑚 compressedtokensand𝑘 = lcm(𝑚,𝑚′) HCAcompressedtokens. 2 𝑚′ For efficient KV cache management of DeepSeek-V4, we design corresponding strategies to overcomethesetwochallenges. StateCacheforSWAandUncompressedTailTokens. Toaddressthefirstobstacle,weadopt analternativecachemanagementmechanism. SinceSWAisdesignedtoenhanceperformance underalimitedKVcachesize,itisreasonabletotreatit,alongwiththeuncompressedtailtokens fromthecompressionbranch,asastate-spacemodel. ThecorrespondingKVcachecanthusbe regardedasasequence-specificstatethatdependssolelyonthecurrentposition. Accordingly, wepre-allocateafixed-andlimited-sizepoolofstatecaches,anddynamicallyassignittoeach sequence. Sparse Attention Kernel Co-Design. Regarding the second obstacle, conventional high- performanceattentionkernelstypicallyassumeafixednumber 𝐵oftokensperblocktooptimize performance, correspondingto 𝐵·𝑚 originaltokensinCSAand 𝐵·𝑚′ inHCA.Throughem- ployingahigh-performancesparse-attentionkernel,differentlayerscanaccommodatevariable tokensperblockwithoutperformancedegradation. Achievingthisrequiresco-designingthe KV cache layout and the sparse attention kernel. For instance, padding blocks to align with cachelinescanimproveperformance. Thus,forCSAwithcompressionratio𝑚andHCAwith ratio 𝑚′, thenumberoforiginaltokensperblockcanbeanymultipleoflcm(𝑚,𝑚′), theleast commonmultipleofthesetwocompressionratios. 3.6.2. On-DiskKVCacheStorage WhenservingDeepSeek-V4,weleverageanon-diskKVcachestoragemechanismtoeliminate repeatedprefillingforshared-prefixrequests. ForthecompressedKVentriesinCSA/HCAand theuncompressedKVentriesinSlidingWindowAttention(SWA),wedesignseparatesolutions forstoragemanagement. 23


For CSA and HCA, we simply store all of the compressed KV entries to the disk. When a request hits a stored prefix, we read and reuse the compressed KV entries corresponding totheprefix,untilthelastcompletecompressionblock. Specially,forprefixtokensinthetail incompleteblock,westillneedtorecomputethemtorestoretheuncompressedKVentries,as uncompressedKVentriesinCSAandHCAarenotstored. FortheSWAKVentries,sincetheyarenotcompressedandexistineverylayer,theirvolume is approximately 8 times larger than the compressed CSA and HCA KV entries. To handle theselargeSWAKVentriesefficiently,weproposeandimplementthreedistinctstrategiesfor managingon-diskSWAKVentries,eachofferingadifferenttrade-offbetweenstorageoverhead andcomputationalredundancy: • Full SWA Caching. This strategy stores the complete SWA KV entries for all tokens, ensuringcomputationalzero-redundancy. Underthisstrategy,theSWAKVentriesofthe hittingprefixcanbereconstructedbyjustreadingtheon-diskcacheofthelast𝑛 tokens win withinthatprefix. Despitecomputationalzero-redundancy,thisstrategyisinefficientfor modernSSD-basedstoragesystems—onlyasmallsubsetofthestoredSWAKVcache will be accessed for each hitting request, which leads to an unbalanced write-intensive accesspattern. • PeriodicCheckpointing. ThisstrategycheckpointsSWAKVentriesofthelast𝑛 tokens win withinevery 𝑝tokens,where 𝑝isatunableparameter. Forahittingprefix,weloadthe mostrecentcheckpointedstate,andthenrecomputetheremainingtailtokens. Through tuning 𝑝,thisstrategyenablesanon-demandtrade-offbetweenstorageandcomputation. • ZeroSWACaching. ThisstrategydoesnotstoreanySWAKVentries. Forahittingprefix, weneedtoperformmorerecomputationtorestoretheSWAKVentries. Tobespecific,in eachattentionlayer,theSWAKVentryofeachtokendependsontheSWAKVentriesof onlythemostrecent 𝑛 tokensfromthepreviouslayer. Therefore, leveragingcached win CSAandHCAKVentries,recomputingthelast𝑛 ·𝐿tokensisenoughtorestorethelast win 𝑛 SWAKVentriesforan 𝐿-layermodel. win Dependingonspecificdeploymentscenarios,weselectthemostsuitablestrategytoachievethe desiredtrade-offbetweenstorageandcomputation. 4. Pre-Training 4.1. DataConstruction Ontopofthepre-trainingdataofDeepSeek-V3,weendeavortoconstructamorediverseand higher-qualitytrainingcorpuswithlongereffectivecontexts. Wecontinuallyrefineourdatacon- structionpipelines. Forweb-sourceddata,weimplementfilteringstrategiestoremovebatched auto-generatedandtemplatedcontent,therebymitigatingtheriskofmodelcollapse(Zhuetal., 2024). Mathematicalandprogrammingcorporastillremaincorecomponentsofourtraining data,andwefurtherenhancethecodingcapabilitiesofDeepSeek-V4seriesbyincorporating agentic data during the mid-training phase. For multilingual data, we build a larger corpus forDeepSeek-V4,improvingitscaptureoflong-tailknowledgeacrossdifferentcultures. For DeepSeek-V4, we place a particular emphasis on long-document data curation, prioritizing scientific papers, technical reports, and other materials that reflect unique academic values. Combiningalltheabove,ourpre-trainingcorpuscomprisesmorethan32Ttokens,containing mathematicalcontents,codes,webpages,longdocuments,andotherhigh-qualitycategories. For pre-training data, we largely follow the same pre-processing strategies of DeepSeek- 24


V3. Fortokenization,ontopoftheDeepSeek-V3tokenizer,weintroduceafewspecialtokens for context construction, and still remain the vocabulary size to be 128K. We also inherit the token-splitting (DeepSeek-AI, 2024) and Fill-in-Middle (FIM) (DeepSeek-AI, 2024) strategies fromDeepSeek-V3. InspiredbyDingetal.(2024),wepackdocumentsfromdifferentsources intoappropriatesequencestominimizesampletruncation. DifferentfromDeepSeek-V3,we employsample-levelattentionmaskingduringpre-training. 4.2. Pre-TrainingSetups 4.2.1. ModelSetups DeepSeek-V4-Flash. WesetthenumberofTransformerlayersto43andthehiddendimension 𝑑 to4096. Forthefirsttwolayers,weusepureslidingwindowattention. Forthesubsequent layers,CSAandHCAareusedinaninterleavedmanner. ForCSA,wesetthecompressionrate 𝑚to4,thenumberofindexerqueryheads𝑛𝐼 to64,theindexerheaddimension𝑐𝐼 to128,and ℎ thenumberofKVentriesselectedforsparseattention(i.e.,attentiontop-k)to512. ForHCA, we set the compression rate 𝑚′ to 128. For both CSA and HCA, we set the number of query heads𝑛 ℎ to64,theheaddimension𝑐to512,andthequerycompressiondimension𝑑 𝑐 to1024. Thenumberofoutputprojectiongroups𝑔 issetto8,andthedimensionofeachintermediate attentionoutput 𝑑 𝑔 issetto1024. Fortheadditionalbranchofslidingwindowattention, the window size 𝑛 is set to 128. We employ MoE layers in all Transformer blocks, but use the win Hashroutingstrategyforthefirst3MoElayers. EachMoElayerconsistsof1sharedexpertand 256routedexperts,wheretheintermediatehiddendimensionofeachexpertis2048. Amongthe routedexperts,6expertswillbeactivatedforeachtoken. Themulti-tokenpredictiondepthis setto1. AsformHC,theexpansionfactor𝑛 issetto4,andthenumberofSinkhorn-Knopp hc iterations𝑡 issetto20. Underthisconfiguration,DeepSeek-V4-Flashcomprises284Btotal max parameters,ofwhich13Bareactivatedforeachtoken. DeepSeek-V4-Pro. WesetthenumberofTransformerlayersto61andthehiddendimension 𝑑 to7168. Forthefirsttwolayers,weuseHCA.Forthesubsequentlayers,CSAandHCAare used in an interleaved manner. For CSA, we set the compression rate 𝑚 to 4, the number of indexerqueryheads𝑛𝐼 to64,theindexerheaddimension𝑐𝐼 to128,andthenumberofKVentries ℎ selected for sparse attention (i.e., attention top-k) to 1024. For HCA, we set the compression rate𝑚′ to128. ForbothCSAandHCA,wesetthenumberofqueryheads𝑛 ℎ to128,thehead dimension 𝑐 to512, andthequerycompressiondimension 𝑑 𝑐 to1536. Thenumberofoutput projectiongroups𝑔issetto16,andthedimensionofeachintermediateattentionoutput𝑑 𝑔 isset to1024. Fortheadditionalbranchofslidingwindowattention,thewindowsize𝑛 issetto win 128. WeemployMoElayersinallTransformerblocks,butusetheHashroutingstrategyforthe first3MoElayers. EachMoElayerconsistsof1sharedexpertand384routedexperts,where theintermediatehiddendimensionofeachexpertis3072. Amongtheroutedexperts,6experts willbeactivatedforeachtoken. Themulti-tokenpredictiondepthissetto1. AsformHC,the expansionfactor𝑛 issetto4,andthenumberofSinkhorn-Knoppiterations𝑡 issetto20. hc max Underthisconfiguration,DeepSeek-V4-Procomprises1.6Ttotalparameters,ofwhich49Bare activatedforeachtoken. 4.2.2. TrainingSetups DeepSeek-V4-Flash. WeemploytheMuonoptimizer(Jordanetal.,2024;Liuetal.,2025)for themajorityofparameters,butusetheAdamWoptimizer(LoshchilovandHutter,2017)forthe 25


embeddingmodule,thepredictionheadmodule,andtheweightsofallRMSNormmodules. For AdamW,wesetitshyper-parametersto 𝛽 =0.9, 𝛽 =0.95,𝜀 =10−20,andweight_decay =0.1. 1 2 ForMuon,wesetthemomentumto0.95andtheweightdecayto0.1,andrescaletheRMSofeach updatematrixto0.18forreutilizationoftheAdamWlearningrate. WetrainDeepSeek-V4-Flash on32Ttokens, andasinDeepSeek-V3, wealsoemployabatchsizeschedulingstrategythat increasesthebatchsize(intokens)fromasmallsizeto75.5Mandthenkeepsitat75.5Mduring mostofthetraining. Thelearningrateislinearlywarmedupinthefirst2000steps,maintained at2.7×10−4 formostofthetraining. Neartheendofthetraining,wefinallydecaythelearning rateto2.7×10−5 followingacosineschedule. Thetrainingstartswithasequencelengthof4K, andwegraduallyextendthetrainingsequencelengthto16K,64K,and1M.Asforthesetupsof sparseattention,wefirstwarmupthemodelwithdenseattentionforthefirst1Ttokens,and introducesparseattentionatthesequencelengthof64Kandkeepsparseattentionduringthe restofthetraining. Whenintroducingattentionsparsity,wefirstsetashortstagetowarmup the lightning indexer in CSA, and then train the model with sparse attention for most of the training. Forauxiliary-loss-freeloadbalancing,wesetthebiasupdatespeedto0.001. Forthe balanceloss,wesetitslossweightto0.0001toavoidextremeimbalancewithinsinglesequences. TheMTPlossweightissetto0.3formostofthetraining,andto0.1uponthestartoflearning ratedecay. DeepSeek-V4-Pro. Except for specific values of hyper-parameters, the training setup of DeepSeek-V4-ProislargelyconsistentwiththatofDeepSeek-V4-Flash. WeemploytheMuonop- timizerforthemajorityofparameters,butusetheAdamWoptimizerfortheembeddingmodule, thepredictionheadmodule,andtheweightsofallRMSNormmodules. Thehyper-parameters ofAdamWandMuonarethesameasthoseofDeepSeek-V4-Flash. WetrainDeepSeek-V4-Pro on 33T tokens, and also employ a batch size scheduling strategy, with the maximum batch size being 94.4M tokens. The learning rate scheduling strategy is largely the same as that of DeepSeek-V4-Flash,butthepeaklearningrateissetto2.0×10−4 andtheendlearningrateisset to2.0×10−5. Thetrainingalsostartswithasequencelengthof4K,andthelengthisgradually extendedto16K,64K,and1M.ComparedwithDeepSeek-V4-Flash, DeepSeek-V4-Prostarts withalongerstageofdenseattention,andthestrategyofintroducingsparseattentionisthe sameasDeepSeek-V4-Flash,followingatwo-stagetrainingmethod. Forauxiliary-loss-freeload balancing,wesetthebiasupdatespeedto0.001. Forthebalanceloss,wesetitslossweightto 0.0001toavoidextremeimbalancewithinsinglesequences. TheMTPlossweightissetto0.3for mostofthetraining,andto0.1uponthestartoflearningratedecay. 4.2.3. MitigatingTrainingInstability Trainingtrillion-parameterMoEmodelspresentssignificantstabilitychallenges,andDeepSeek- V4 series are no exception. We encountered notable instability challenges during training. Whilesimplerollbackscouldtemporarilyrestorethetrainingstate,theyprovedinadequateasa long-termsolutionbecausetheydonotpreventtherecurrenceoflossspikes. Empirically,we identifiedthattheoccurrenceofspikesisconsistentlytiedtooutliersintheMoElayers,andthe routingmechanismitselfappearstoexacerbatetheemergenceoftheseoutliers. Therefore,we soughttotacklethisissuefromtwodimensions: breakingtheviciouscycleinducedbyrouting, anddirectlysuppressinganomalousvalues. Fortunately,wediscoveredtwopracticaltechniques thateffectivelymaintaintrainingstability. Althoughacomprehensivetheoreticalunderstanding oftheirunderlyingmechanismsremainsanopenquestionfornow,wearesharingthemopenly tofosterfurtherexplorationbythecommunity. 26


AnticipatoryRouting. Wefoundthatdecouplingthesynchronousupdatesofthebackbone networkandtheroutingnetworksignificantlyimprovestrainingstability. Consequently,atstep 𝑡,weusethecurrentnetworkparameters𝜃 𝑡 forfeaturecomputation,buttheroutingindicesare computedandappliedusingthehistoricalnetworkparameters𝜃 𝑡−Δ𝑡. Inpractice,tocircumvent the overhead of loading model parameters twice, we fetch the data for step 𝑡 in advance at step𝑡−Δ𝑡. We"anticipatorily"computeandcachetheroutingindicestobeusedlateratstep 𝑡,whichiswhywenamethisapproachAnticipatoryRouting. Wealsoheavilyoptimizedthis at the infrastructure level. First, given that pre-computing the routing indices only requires asingleforwardpassoverthedata,wecarefullyorchestratedthepipelineexecutionandthe overlappingofcomputationwithExpertParallelism(EP)communication,successfullybounding theadditionalwall-clocktimeoverheadofAnticipatoryRoutingtoapproximately20%. Second, weintroducedanautomaticdetectionmechanismthattriggersashortrollbackandactivates AnticipatoryRoutingexclusivelywhenalossspikeoccurs;afteroperatinginthismodefora certain period, the system reverts to standard training. Ultimately, this dynamic application allowsustoavertlossspikeswithnegligibleoveralladditionaltrainingoverhead,allwithout compromisingmodelperformance. SwiGLUClamping. Inpreviousliterature(Belloetal.,2017;Riviereetal.,2024), clamping hasbeenexplicitlyutilizedtoconstrainnumericalranges,therebyenhancingtrainingstability. Inouractualtrainingruns,weempiricallyfoundthatapplyingSwiGLUclamping(OpenAI, 2025) effectively eliminates outliers and substantially aids in stabilizing the training process, withoutcompromisingperformance. ThroughoutthetrainingofbothDeepSeek-V4-Flashand DeepSeek-V4-Pro,weclampedthelinearcomponentofSwiGLUtotherangeof [−10,10],while cappingtheupperboundofthegatecomponentat10. 4.3. Evaluations 4.3.1. EvaluationBenchmarks Fortheevaluationofthebasemodels,weconsiderbenchmarksspanningfourkeydimensions: worldknowledge,languageunderstandingandreasoning,codingandmathematics,andlong- contextprocessing. WorldknowledgebenchmarksincludeAGIEval(Zhongetal.,2023),C-Eval(Huangetal., 2023), CMMLU (Li et al., 2023) MMLU (Hendrycks et al., 2020), MMLU-Redux (Gema et al., 2024), MMLU-Pro (Wang et al., 2024b), MMMLU (OpenAI, 2024a), MultiLoKo (Hupkes and Bogoychev,2025),Simple-QAverified(Haasetal.,2025),SuperGPQA(Duetal.,2025),FACTS Parametric(Chengetal.,2025),andTriviaQA(Joshietal.,2017). LanguageunderstandingandreasoningbenchmarksincludeBigBenchHard(BBH)(Suzgun etal.,2022),DROP(Duaetal.,2019),HellaSwag(Zellersetal.,2019),CLUEWSC(Xuetal.,2020), andWinoGrande(Sakaguchietal.,2019). Coding and mathematical benchmarks include BigCodeBench (Zhuo et al., 2025), Hu- manEval(Chenetal.,2021),GSM8K(Cobbeetal.,2021),MATH(Hendrycksetal.,2021),MGSM (Shietal.,2023),andCMath(Weietal.,2023). LongcontextbenchmarksincludeLongBench-V2(Baietal.,2025b). 27


Table1 | ComparisonamongDeepSeek-V3.2-Base,DeepSeek-V4-Flash-Base,andDeepSeek-V4- Pro-Base. Allmodelsareevaluatedinourinternalframeworkandsharethesameevaluation setting. Scoreswithagapnotexceeding0.3areconsideredtobeatthesamelevel. Thehighest scoreineachrowisinboldfont,andthesecondisunderlined. DeepSeek-V3.2 DeepSeek-V4-Flash DeepSeek-V4-Pro Benchmark(Metric) #Shots Base Base Base Architecture - MoE MoE MoE

ActivatedParams - 37B 13B 49B

TotalParams - 671B 284B 1.6T

AGIEval(EM) 0-shot 80.1 82.6 83.1 MMLU(EM) 5-shot 87.8 88.7 90.1 MMLU-Redux(EM) 5-shot 87.5 89.4 90.8 MMLU-Pro(EM) 5-shot 65.5 68.3 73.5 MMMLU(EM) 5-shot 87.9 88.8 90.3 C-Eval(EM) 5-shot 90.4 92.1 93.1 WorldKnowl. CMMLU(EM) 5-shot 88.9 90.4 90.8 MultiLoKo(EM) 5-shot 38.7 42.2 51.1 Simple-QAverified(EM) 25-shot 28.3 30.1 55.2 SuperGPQA(EM) 5-shot 45.0 46.5 53.9 FACTSParametric(EM) 25-shot 27.1 33.9 62.6 TriviaQA(EM) 5-shot 83.3 82.8 85.6 BBH(EM) 3-shot 87.6 86.9 87.5 DROP(F1) 1-shot 88.2 88.6 88.7 Lang.&Reas. HellaSwag(EM) 0-shot 86.4 85.7 88.0 WinoGrande(EM) 0-shot 78.9 79.5 81.5 CLUEWSC(EM) 5-shot 83.5 82.2 85.2 BigCodeBench(Pass@1) 3-shot 63.9 56.8 59.2 HumanEval(Pass@1) 0-shot 62.8 69.5 76.8 GSM8K(EM) 8-shot 91.1 90.8 92.6 Code&Math MATH(EM) 4-shot 60.5 57.4 64.5 MGSM(EM) 8-shot 81.3 85.7 84.4 CMath(EM) 3-shot 92.6 93.6 90.9 LongContext LongBench-V2(EM) 1-shot 40.2 44.7 51.5 4.3.2. EvaluationResults InTable1,weprovideadetailedcomparisonofthebasemodelsforDeepSeek-V3.2,DeepSeek- V4-Flash,andDeepSeek-V4-Pro,allevaluatedunderaunifiedinternalframeworkwithstrictly consistentsettings. Comparing DeepSeek-V4-Flash-Base with DeepSeek-V3.2-Base reveals a compelling ef- ficiency story. Despite utilizing a substantially smaller number of both activated and total parameters,DeepSeek-V4-Flash-BaseoutperformsDeepSeek-V3.2-Baseacrossawidearrayof benchmarks. Thisadvantageisespeciallyevidentinworldknowledgetasksandchallenging long-contextscenarios. Theseresultsunderscorethatarchitecturalimprovements,refineddata quality,andtrainingoptimizationsinDeepSeek-V4-Flash-Baseyieldsuperiorperformanceeven withamorecompactparameterbudget,effectivelysurpassingthelargerDeepSeek-V3.2-Base onthemajorityofevaluations. Furthermore, DeepSeek-V4-Pro-Base demonstrates a further, decisive leap in capability, establishingnear-universaldominanceoverbothDeepSeek-V3.2-BaseandDeepSeek-V4-Flash- Base. With improvements across almost all categories, DeepSeek-V4-Pro-Base reaches new 28


performance highs among DeepSeek base models on the most demanding benchmarks. On knowledge-intensiveevaluations,itdeliversdramaticgains,whilealsosubstantiallyadvancing long-contextunderstanding. Onmostreasoningandcodebenchmarks,DeepSeek-V4-Pro-Base alsoexceedsbothpreviousmodels. ThiscomprehensiveupliftconfirmsDeepSeek-V4-Pro-Base asthestrongestfoundationmodelintheDeepSeekseries,outperformingitspredecessorsacross thespectrumofknowledge,reasoning,coding,andlong-contextcapabilities. 5. Post-Training 5.1. Post-TrainingPipeline Followingpre-training,weconductedapost-trainingphasetoyieldthefinalmodelsofDeepSeek- V4 series. Although the training pipeline largely mirrored that of DeepSeek-V3.2, a critical methodological substitution was made: the mixed Reinforcement Learning (RL) stage was entirelyreplacedbyOn-PolicyDistillation(OPD). 5.1.1. SpecialistTraining ThedevelopmentofdomainspecialistswasconductedbyadaptingtheDeepSeek-V3.2training pipeline. Specifically, each model was sequentially optimized through an initial fine-tuning phaseandsubsequentReinforcementLearning(RL)guidedbydomain-specificpromptsandre- wardsignals. FortheRLstage,weimplementedtheGroupRelativePolicyOptimization(GRPO) algorithm,maintaininghyper-parameterscloselyalignedwithourpriorresearch(DeepSeek-AI, 2025;DeepSeek-AI,2025). Reasoning Efforts. It is widely recognized that a model’s performance on reasoning tasks isfundamentallygovernedbythecomputationaleffortexpended. Consequently,wetrained distinctspecialistmodelsunderdivergentRLconfigurationstofacilitatethedevelopmentof modelsoptimizedforvaryingreasoningcapacities. AsdetailedinTable2,DeepSeek-V4-Proand DeepSeek-V4-Flashbothsupportthreespecificreasoningeffortmodes. Foreachmode,weapply distinct length penalties and context windows during RL training, which results in varying output token lengths for reasoning. To integrate these distinct reasoning modes, we utilize specializedresponseformatsdemarcatedbytheandtokens. Furthermore, for the "Think Max" mode, we prepend a specific instruction to the beginning of the system prompttoguidethemodel’sreasoningprocess,asshowninTable3. GenerativeRewardModel. Typically,easy-to-verifytaskscanbeeffectivelyoptimizedusing simplerule-basedverifiersortestcases. Incontrast,hard-to-verifytaskstraditionallyrelyon ReinforcementLearningfromHumanFeedback(RLHF),whichnecessitatesextensivehuman annotation to train a scalar reward model. In the post-training phase of DeepSeek-V4 series, however,wedispensewiththeseconventionalscalar-basedrewardmodels. Instead,toaddress hard-to-verifytasks,wecuraterubric-guidedRLdataandemployaGenerativeRewardModel (GRM)toevaluatepolicytrajectories. Crucially,weapplyRLoptimizationdirectlytotheGRM itself. In this paradigm, the actor network natively functions as the GRM, enabling the joint optimizationofthemodel’sevaluative(judging)proficiencyalongsideitsstandardgenerative capabilities. Byunifyingtheseroles,themodel’sinternalreasoningcapabilitiesareinherently fusedintoitsevaluativeprocess,resultinginhighlyrobustscoring. Furthermore,thisapproach achievessuperiorperformancewithonlyaminimalsetofdiversehumanannotations,asthe 29


Table2 | Comparisonofthreereasoningmodes Reasoning Characteristics TypicalUseCases ResponseFormat Mode Non-think Fast, intuitive re- Routine daily tasks, summary sponses based on emergencyreactions, habits or simple low-riskdecisions. rules. ThinkHigh Conscious logical Complex problem- thinking analysis, slower but solving, planning, tokens moreaccurate. medium-risk deci- summary sions. ThinkMax Pushreasoningtoits Exploringthebound- 1. A special system fullest extent. Slow aryofmodelreason- promptatthebegin- butpowerful. ingcapability. ning. 2. thinking tokens summary Table3 | Instructioninjectedintothesystempromptforthe"ThinkMax"mode. InjectedInstruction ReasoningEffort: Absolutemaximumwithnoshortcutspermitted. You MUST be very thorough in your thinking and comprehensively decompose the problemtoresolvetherootcause,rigorouslystress-testingyourlogicagainstallpotential paths,edgecases,andadversarialscenarios. Explicitlywriteoutyourentiredeliberationprocess, documentingeveryintermediate step,consideredalternative,andrejectedhypothesistoensureabsolutelynoassumption isleftunchecked. modelleveragesitsownlogictogeneralizeacrosscomplextasks. Tool-Call Schema and Special Token. Consistent with our previous version, we utilize a dedicated tag to delineate the reasoning path. In DeepSeek-V4 series, we introduceanewtool-callschemathatemploysaspecial"|DSML|"tokenandutilizesanXML- basedformatfortoolinvocations,asdemonstratedinTable4. Ourexperimentsdemonstratethat theXMLformateffectivelymitigatesescapingfailuresandreducestool-callerrors,providinga morerobustinterfaceformodel-toolinteractions. InterleavedThinking. DeepSeek-V3.2introducedacontextmanagementstrategythatretains reasoningtracesacrosstool-resultroundsbutdiscardsthemuponthearrivalofnewusermes- sages. Whileeffective,thisstillcausedunnecessarytokenwasteincomplexagenticworkflows — each new user turn would flush all accumulated reasoning content, forcing the model to reconstructitsproblem-solvingstatefromscratch. Leveragingtheexpanded1M-tokencontext 30


Table4 | Tool-callschemaforDeepSeek-V4series. ToolCallSchema

Tools

You have access to a set of tools to help answer the user’s question. You can invoke tools by writing a "<|DSML|tool_calls>" block like the following: <|DSML|tool_calls> <|DSML|invoke name="$TOOL_NAME"> <|DSML|parameter name="$PARAMETER_NAME" string="true|false">$PARAMETER_VALUE </|DSML|parameter> ... </|DSML|invoke> <|DSML|invoke name="$TOOL_NAME2"> ... </|DSML|invoke> </|DSML|tool_calls> String parameters should be specified as is and set ‘string="true"‘. For all other types (numbers, booleans, arrays, objects), pass the value in JSON format and set ‘string="false"‘. If thinking_mode is enabled (triggered by ), you MUST output your complete reasoning inside ... BEFORE any tool calls or final response. Otherwise, output directly after with tool calls or final response.

Available Tool Schemas

{Tool Definition...} You MUST strictly follow the above definedtool name and parameter schemas to invoke tool calls. windowofDeepSeek-V4series,wefurtherrefinethismechanismtomaximizetheeffectiveness ofinterleavedthinkinginagenticenvironments: • Tool-Calling Scenarios. As illustrated in Figure 7(a), all reasoning content is fully pre- served throughout the entire conversation. Unlike DeepSeek-V3.2, which discarded thinkingtracesuponeachnewuserturn,DeepSeek-V4seriesretainthecompletereason- inghistoryacrossallrounds,includingacrossusermessageboundaries. Thisallowsthe modeltomaintainacoherent,cumulativechainofthoughtoverlong-horizonagenttasks. • GeneralConversationalScenarios. AsillustratedinFigure7(b),theoriginalstrategyis preserved: reasoningcontentfrompreviousturnsisdiscardedwhenanewusermessage arrives,keepingthecontextconciseforsettingswherepersistentreasoningtracesprovide limitedbenefit. AswithDeepSeek-V3.2,agentframeworksthatsimulatetoolinteractionsviausermessages(e.g., Terminus)maynottriggerthetool-callingcontextpathandthusmaynotbenefitfromenhanced reasoningpersistence. Wecontinuetorecommendnon-thinkmodelsforsucharchitectures. 31


a) Thinking with tools b) Thinking without tools Figure7 | ThinkingmanagementofDeepSeek-V4series. QuickInstruction. Inchatbotscenarios,anumberofauxiliarytasks(e.g.,determiningwhether totriggerawebsearch,intentrecognition,etc.) mustbeexecutedbeforegeneratingtheresponse. Conventionally,thesetasksarehandledbyaseparatesmallmodel,requiringredundantprefill- ingsinceitcannotreusetheexistingKVcache. Toovercomethislimitation,weintroduceQuick Instruction. Weappendasetofdedicatedspecialtokensdirectlytotheinputsequence,where eachtokencorrespondstoaspecificauxiliarytask. Bydirectlyreusingthealready-computed KVcache,thismechanismcompletelyavoidsredundantprefillingandallowscertaintasks,such asgeneratingsearchqueriesanddeterminingauthorityanddomain,tobeexecutedinparallel. Consequently,thisapproachsignificantlyreducestheuser-perceivedtime-to-first-token(TTFT) andeliminatestheengineeringoverheadofmaintaininganditeratinganextrasmallmodel. The supportedQuickInstructiontokensaresummarizedinTable5. 5.1.2. On-PolicyDistillation Aftertrainingmultipledomain-specificexpertsviaspecializedfine-tuningandreinforcement learning,weemploymulti-teacherOn-PolicyDistillation(OPD)astheprimarytechniquefor mergingexpertcapabilitiesintothefinalmodel. OPDhasemergedasaneffectivepost-training paradigm for efficiently transferring the knowledge and capabilities of domain experts to a single,unifiedmodel. Thisisachievedbyhavingthestudentlearnfromtheoutputdistributions ofteachermodelsonitsowngeneratedtrajectories. Formally,givenasetof 𝑁 expertmodels 32


Table5 | QuickInstructionspecialtokensforauxiliarytasks. SpecialToken Description Format <|action|> Determineswhethertheuser ...<|User|>{prompt}<|Assistant|> promptrequiresawebsearch <|action|> orcanbeanswereddirectly. <|title|> Generatesaconciseconversa- ...<|Assistant|>{response} tion title after the first assis- <|end_of_sentence|><|title|> tantresponse. <|query|> Generatessearchqueriesfor ...<|User|>{prompt}<|query|> theuserprompt. <|authority|> Classifies the user prompt’s ...<|User|>{prompt}<|authority|> demandforsourceauthorita- tiveness. <|domain|> Identifies the domain of the ...<|User|>{prompt}<|domain|> userprompt. <|extracted_url|>Determines whether each ...<|User|>{prompt} <|read_url|> URL in the user prompt <|extracted_url|>{url} shouldbefetchedandread. <|read_url|> {𝜋 𝐸 1 ,𝜋 𝐸 2 ,...,𝜋 𝐸𝑁 },theOPDobjectivefunctionisdefinedas: 𝑁 L OPD (𝜃) = ∑︁ 𝑤 𝑖 ·D KL (cid:0)𝜋 𝜃 ∥ 𝜋 𝐸𝑖 (cid:1) . (29) 𝑖=1 Inthisformulation,𝑤 𝑖 representstheassignedweightforeachexpert,typicallydeterminedby the relative importance of the expert. Computing the reverse KL loss D KL (cid:0)𝜋 𝜃 ∥ 𝜋 𝐸𝑖 (cid:1) requires samplingtrainingtrajectoriesfromthestudent𝜋 𝜃 tomaintainon-policylearning. Theunderly- inglogicensuresthattheunifiedpolicy𝜋 𝜃selectivelylearnsfromthespecializedexpertrelevant tothecurrenttaskcontext(e.g.,aligningwiththemathematicsexpertformathreasoningtasks andthecodingexpertforprogrammingtasks). Throughthismechanism,theknowledgefrom physicallydistinctexpertweightsisconsolidatedintoaunifiedparameterspacevialogits-level alignment, practically circumventing the performance degradation often encountered in tra- ditionalweight-mergingormixedRLtechniques. Inthisstage,morethantenteachermodels coveringvariousdomainsareemployedtodistillasinglestudentmodel. InhandlingtheaboveOPDobjective,priorworksusuallysimplifythefull-vocabularyKL lossintoatoken-levelKLestimateateachtokenposition,andreuseRLframeworkbyreplac- ingsg (cid:2) log 𝜋𝐸𝑖 (𝑦𝑡|𝑥,𝑦<𝑡)(cid:3) (sgrepresentsthestopgradientoperation)astheper-tokenadvantage 𝜋 𝜃(𝑦𝑡|𝑥,𝑦<𝑡) estimate in the policy loss calculation. Although this approach is resource-efficient, it leads to high variance in gradient estimation and often causes training instability. Therefore, we adoptfull-vocabularylogitdistillationinourOPD.Preservingthecompletelogitdistributionin calculatingreverseKLlossyieldsmorestablegradientestimatesandensuresfaithfuldistillation oftheteachers’knowledge. Inthefollowingsubsection,wedescribetheengineeringeffortsthat makefull-vocabularyOPDfeasibleatscale. 33


5.2. RLandOPDInfrastructures Ourpost-traininginfrastructureisbuiltuponthescalableframeworkdevelopedforDeepSeek- V3.2. Specifically,weintegratethesamedistributedtrainingstackdescribedinSection3.5and the rollout engine introduced earlier for efficient auto-regressive sampling. Building on this foundation, we introduce the following principal enhancements in the present work. These designsenableefficientexecutionofultra-long-contextRLandOPDmergingtasksinvolving overtendistinctteachermodels,therebysubstantiallyacceleratingtheiterationcycleformodel releases. 5.2.1. FP4QuantizationIntegration WeapplyFP4(MXFP4)quantizationtoacceleratebothrolloutsandallinference-onlyforward passes,includingthoseofteacherandreferencemodels,therebyreducingmemorytrafficand sampling latency. As detailed in Section 3.4, we directly use native FP4 weights during the rollout and inference phases. For training steps, FP4 quantization is simulated via a lossless FP4-to-FP8dequantizationstep,allowingseamlessreuseoftheexistingFP8mixed-precision frameworkwithFP32masterweightsandrequiringnomodificationtothebackwardpipeline. 5.2.2. EfficientTeacherSchedulingforFull-VocabularyOPD Our framework supports full-vocabulary On-Policy Distillation (OPD) with an effectively unboundednumberofteachers,eachpotentiallycomprisingtrillionsofparameters. Toenable this, all teacher weights are offloaded to a centralized distributed storage and are loaded on demandduringtheteacherforwardpasswithZeRO-likeparametershardingtoalleviateboth I/O and DRAM pressure. Furthermore, naively materializing logits for a vocabulary size |𝑉| > 100k across all teachers is prohibitive, even when spooled to disk. We address this by caching only the last-layer teacher hidden states in a centralized buffer during the forward pass. Attrainingtime,thesecachedstatesareretrievedandpassedthroughthecorresponding predictionheadmoduletoreconstructthefulllogitsonthefly. Thisdesignincursnegligible recomputationoverheadwhilecompletelycircumventingthememoryburdenassociatedwith explicitlogitsmaterialization. TomitigatetheGPUmemoryfootprintoftheteacherprediction head,weordertrainingsamplesbyteacherindexduringdatadispatching. Thisarrangement ensures that each distinct teacher head is loaded only once per mini-batch and that at most oneteacherheadresidesindevicememoryatanygiventime. Allparametersandhiddenstate loading/offloadingoperationsproceedasynchronouslyinthebackground,withoutblocking computationonthecriticalpath. Finally,theexactKLdivergencesbetweenteacherandstudent logitsarecomputedusingaspecializedTileLangkernel,whichacceleratesthecomputationand curtailsdynamicmemoryallocation. 5.2.3. PreemptibleandFault-TolerantRolloutService TomaximizeGPUresourceutilizationwhileenablingrapidhardwareprovisioningforhigh- prioritytasks,ourGPUclusteremploysacluster-widepreemptivetaskscheduler,whereany runningtaskmaybepreemptedatanytime. Also,hardwarefailuresareprevalentinlarge-scale GPU clusters. To this end, we implement a preemptible and fault-tolerant LLM generation serviceforRL/OPDrollout. Specifically,weimplementatoken-granularWrite-AheadLog(WAL)foreachgeneration request. Wheneveranewtokenisgeneratedforarequest,weimmediatelyappendittothat request’sWAL.Duringpreemption,wepausetheinferenceengineandsavetheKVcacheof 34


unfinished requests. Upon resumption, we use the persisted WALs and saved KV cache to continuedecoding. Evenwhenafatalhardwareerroroccurs,wecanre-runtheprefillphase usingthepersistedtokensinWALtoreconstructtheKVcache. Importantly,itismathematicallyincorrecttoregenerateunfinishedrequestsfromscratch, asthisintroduceslengthbias. Becauseshorterresponsesaremorelikelytosurviveinterrup- tion,regeneratingfromscratchmakesthemodelmorepronetoproducingshortersequences whenever an interruption occurs. If the inference stack is batch-invariant and deterministic, this correctness issue could also be addressed by regenerating with a consistent seed for the pseudorandomnumbergeneratorusedinthesampler. However,thisapproachstillincursthe extracostofre-runningthedecodingphase,makingitfarlessefficientthanourtoken-granular WALmethod. 5.2.4. ScalingRLFrameworkforMillion-TokenContext We introduce targeted optimizations for efficient RL and OPD on million-token sequences. Duringtherolloutphase,weadoptapreemptibleandfault-tolerantrolloutservice,detailedin Section5.2.3. Fortheinferenceandtrainingphase,wedecomposetherolloutdataformatinto lightweightmetadataandheavyper-tokenfields. Duringdatadispatching,themetadataforthe entirerolloutdatacanbeloadedtoperformglobalshufflingandpackinglayoutcomputation. Heavy per-token fields are loaded via a shared-memory data loader to eliminate intra-node dataredundancyandarereleasedimmediatelyuponconsumptionatthemini-batchgranularity, substantiallyreducingbothCPUandGPUmemorypressure. Thenumberofon-devicemini- batchesisdynamicallydeterminedbasedonworkload,allowinganefficienttrade-offbetween computationalthroughputandI/Ooverlap. 5.2.5. SandboxInfrastructureforAgenticAI To meet the diverse execution demands of agentic AI during post-training and evaluation, we build a production-grade sandbox platform, DeepSeek Elastic Compute (DSec). DSec comprisesthreeRustcomponents—theAPIgateway(Apiserver),per-hostagent(Edge),and theclustermonitor(Watcher)—thatareinterconnectedbyacustomRPCprotocolandscale horizontallyatopthe3FSdistributedfilesystem(DeepSeek-AI,2025). Inproduction,asingle DSecclustermanageshundredsofthousandsofconcurrentsandboxinstances. The design of DSec is motivated by four observations: (1) agentic workloads are highly heterogeneous,spanninglightweightfunctioncallstofullsoftware-engineeringpipelineswith diverseOSandsecurityrequirements;(2)environmentimagesarenumerousandlarge,yetmust loadquicklyandsupportiterativecustomization;(3)high-densitydeploymentdemandsefficient CPUandmemoryutilization;(4)sandboxlifecyclesmustcoordinatewithGPUtrainingsched- ules,includingpreemptionandcheckpoint-basedresumption. Basedontheseobservations,we elaborateonthefourcoredesignsofDSecindividuallyinthefollowing. Four Execution Substrates Behind One Unified Interface. DSec exposes a single Python SDK (libdsec) that abstracts four execution substrates. Function Call dispatches stateless invocationstoapre-warmedcontainerpool,eliminatingcold-startoverhead. Containerisfully Docker-compatible and leverages EROFS (Gao et al., 2019) on-demand loading for efficient imageassembly. microVM,builtonFirecracker(Agacheetal.,2020),addsVM-levelisolationfor security-sensitive,high-densitydeployments. fullVM,builtonQEMU(Bellard,2005),supports arbitraryguestoperatingsystems. AllfourshareacommonAPIsurface—commandexecution, 35


filetransfer,andTTYaccess—andswitchingbetweenthemrequiresonlyaparameterchange. FastImageLoadingviaLayeredStorage. DSecreconcilesfaststartupwithalargeandgrowing corpusofenvironmentimagesthroughlayered,on-demandloading. Forcontainers,baseimages andfilesystemcommitsarestoredas3FS-backedreadonlyEROFSlayersmounteddirectlyinto overlaylowerdirs. Wekeepfilemetadatareadilyavailableonthelocaldiskatmounttime; meanwhile, data blocks are fetched from 3FS upon request. For microVMs, DSec uses the overlaybd (Li et al., 2020) disk format: the read-only base layer resides on 3FS for cross- instancesharing,whilewritesgotoalocalcopy-on-writelayer. Suchsnapshotsarechainable, facilitatingefficientversioningandmillisecond-scaleresumption. DensityOptimizationsUnderMassiveConcurrency. Toaccommodatehundredsofthousands of sandboxes per cluster, DSec tackles two resource bottlenecks. First, it mitigates duplicate page-cachefootprintsinvirtualizedenvironmentsandappliesmemoryreclamationtoenable safeovercommitment. Second,italleviatesspinlockcontentioninthecontainerruntimeand therefore,reducesper-sandboxCPUoverhead,significantlyincreasingper-hostpackingdensity. TrajectoryLoggingandPreemption-SafeResumption. DSecmaintainsagloballyordered trajectorylogforeachsandbox,persistentlyrecordingeverycommandinvocationanditsresults. The trajectory serves three purposes: (1) client fast-forwarding — when a training task is preempted,sandboxresourcesareretainednonetheless;uponresumption,DSecreplayscached resultsforpreviouslycompletedcommands,acceleratingtaskrecoverywhilstalsopreventing errors from re-execution of non-idempotent operations; (2) fine-grained provenance — the originandcorrespondingoutcomesofeachstatechangearetraceable;(3)deterministicreplay —anyhistoricalsessioncanbefaithfullyreproducedfromitstrajectory. 5.3. StandardBenchmarkEvaluation 5.3.1. EvaluationSetup KnowledgeandReasoning. KnowledgeandreasoningdatasetsincludeMMLU-Pro(Wang etal.,2024b),GPQA(Reinetal.,2023),HumanLastExam(Phanetal.,2025),Simple-QAVeri- fied(Haasetal.,2025),Chinese-SimpleQA(Heetal.,2024),LiveCodeBench-v6(Jainetal.,2024), CodeForces(InternalBenchmark),HMMT2026Feb,Apex(Balunovic´ etal.,2025),ApexShort- list(Balunovic´ etal.,2025),IMOAnswerBench(Luongetal.,2025),andPutnamBench(Tsoukalas etal.,2024). Forcode,weevaluateDeepSeek-V4seriesonLiveCodeBench-v6andaninternalCodeforces benchmark. For Codeforces, we collect 14 Codeforces Division 1 contests comprising 114 problems(May2025-November2025). TheEloratingiscomputedasfollows. Foreachcontest, wegenerate32candidatesolutionsperproblem. Foreachproblemindependently,wesample 10 of these solutions without replacement and arrange them in a random order to form the submission sequence. Each submission is judged against a test suite constructed by domain experts. ThescoreforasolvedproblemfollowsthepenaltyschemeofOpenAI(2025): themodel receivesthemedianscoreofhumanparticipantswhosolvedthesameproblemwiththesame numberofpriorfailedattempts. Thisyieldsatotalcontestscoreforeachsampledsubmission sequence,whichisthenconvertedintoacontestrankandsubsequentlyintoanestimatedrating viathestandardCodeforcesratingsystem. Thecontest-levelexpectedratingisdefinedasthe 36


expectationofthis estimated ratingoverallpossiblerandom selectionsandorderingsof the 10 submissions per problem. The model’s overall rating is the average of these contest-level expectedratingsacrossall14contests. Forreasoningandknowledgetasks,wesetthetemperatureto1.0andthecontextwindowto 8K,128K,and384KtokensfortheNon-think,High,andMaxmodes,respectively. Formathtasks (e.g.,HMMT,IMOAnswerBench,Apex,andHLE),weevaluateusingthefollowingtemplate: "{question}\nPlease reason step by step, and put your final answer within \boxed{}."ForDeepSeek-V4-Pro-Maxonmathtasks,weusethefollowingtemplatetoelicit deeper reasoning: "Solve the following problem. The problem may ask you to prove a statement, or ask for an answer. If finding an answer is required, you should come up with the answer, and your final solution should also be a rigorous proof of that answer being valid.\n\n{question}". For formal math tasks, we evaluate in an agentic setting on Lean v4.28.0-rc1 (Moura and Ullrich,2021),withaccesstotheLeancompilerandasemantictacticsearchengine,running up to 500 tool calls with max reasoning effort. In addition, we evaluate a more compute- intensivepipelineinwhichcandidatenatural-languagesolutionsarefirstgeneratedandfiltered byself-verification(Shaoetal.,2025),andtheretainedsolutionsarethenprovidedasguidance to a formal agent for proving the corresponding Lean statement. This design uses informal reasoningtoimproveexplorationwhilepreservingstrictcorrectnessthroughformalverification. A submission is counted as correct only if the strict verifier Comparator accepts it for both settings. WehaveleftsomeentriesblankforK2.6andGLM-5.1,astheirAPIsweretoobusytoreturn responsestoourqueries. 1M-TokenContext. SinceDeepSeek-V4seriessupports1M-tokencontexts,weevaluatemodel performance in a long context scenario by selecting OpenAI MRCR (OpenAI, 2024b) and CorpusQA(Luetal.,2026)asthebenchmarks. Were-evaluateClaudeOpus4.6andGemini3.1 Proonthesetaskswiththegoalofstandardizingtheconfigurationacrossallmodels. Wedid notevaluateGPT-5.4becauseitsAPIfailedtorespondtoalargeportionofourqueries. Agent. AgentdatasetsincludeTerminalBench2.0(Merrilletal.,2026),SWE-Verified(OpenAI, 2024e),SWEMultilingual(Yangetal.,2025),SWE-Pro(Dengetal.,2025),BrowseComp(Wei etal.,2025),thepublicevaluationsetofMCPAtlas(Bandietal.,2026),GDPval-AA(AA,2025; Patwardhanetal.,2025),andTool-Decathlon(Lietal.,2025). Forcodeagenttasks(SWE-Verified,Terminal-Bench,SWE-Pro,SWEMultilingual),weeval- uateDeepSeek-V4seriesusinganinternallydevelopedevaluationframework. Thisframework provides a minimal set of tools — a bash tool and a file-edit tool. The maximum number of interactionstepsissetto500,andthemaximumcontextlengthissetto512Ktokens. Regarding Terminal-Bench2.0,weacknowledgetheenvironment-relatedissuesnotedbyGLM-5.1. Never- theless,wereportourperformanceontheoriginalTerminal-Bench2.0datasetforconsistency. OntheTerminal-Bench2.0Verifiedsubset,DeepSeek-V4-Proachievesascoreofapproximately 72.0. Forsearchagenttasks(BrowseComp,HLEw/tool),wealsouseanin-househarnesswith websearchandPythontool,andsetmaximuminteractionstepsto500andthemaximumcontext length to 512K tokens. For BrowseComp, we use the same discard-all context management strategyasDeepSeek-V3.2(DeepSeek-AI,2025). 37


5.3.2. EvaluationResults Table6 | ComparisonbetweenDeepSeek-V4-Pro-Maxandclosed/opensourcemodels. "Max", "xHigh", and "High" denote reasoning effort. The best results are highlighted in bold; the second-bestresultsareunderlined. Opus-4.6GPT-5.4Gemini-3.1-Pro K2.6 GLM-5.1 DS-V4-Pro Benchmark(Metric) Max xHigh High ThinkingThinking Max gninosaeR&egdelwonK MMLU-Pro(EM) 89.1 87.5 91.0 87.1 86.0 87.5 SimpleQA-Verified(Pass@1) 46.2 45.3 75.6 36.9 38.1 57.9 Chinese-SimpleQA(Pass@1) 76.4 76.8 85.9 75.9 75.0 84.4 GPQADiamond(Pass@1) 91.3 93.0 94.3 90.5 86.2 90.1 HLE(Pass@1) 40.0 39.8 44.4 36.4 34.7 37.7 LiveCodeBench(Pass@1) 88.8 - 91.7 89.6 - 93.5 Codeforces(Rating) - 3168 3052 - - 3206 HMMT2026Feb(Pass@1) 96.2 97.7 94.7 92.7 89.4 95.2 IMOAnswerBench(Pass@1) 75.3 91.4 81.0 86.0 83.8 89.8 Apex(Pass@1) 34.5 54.1 60.9 24.0 11.5 38.3 ApexShortlist(Pass@1) 85.9 78.1 89.1 75.5 72.4 90.2 gnoL MRCR1M(MMR) 92.9 - 76.3 - - 83.5 CorpusQA1M(ACC) 71.7 - 53.8 - - 62.0 citnegA TerminalBench2.0(Acc) 65.4 75.1 68.5 66.7 63.5 67.9 SWEVerified(Resolved) 80.8 - 80.6 80.2 - 80.6 SWEPro(Resolved) 57.3 57.7 54.2 58.6 58.4 55.4 SWEMultilingual(Resolved) 77.5 - - 76.7 73.3 76.2 BrowseComp(Pass@1) 83.7 82.7 85.9 83.2 79.3 83.4 HLEw/tools(Pass@1) 53.1 52.0 51.6 54.0 50.4 48.2 GDPval-AA(Elo) 1619 1674 1314 1482 1535 1554 MCPAtlasPublic(Pass@1) 73.8 67.2 69.2 66.6 71.8 73.6 Toolathlon(Pass@1) 47.2 54.6 48.8 50.0 40.7 51.8 ThecomparisonofDeepSeek-V4-Pro-Maxandotherclosed/opensourcemodelsispresented inTable6. Also,weevaluatedifferentmodesofDeepSeek-V4-FlashandDeepSeek-V4-Proand showtheresultsinTable7. Knowledge. Intheevaluationofgeneralworldknowledge,DeepSeek-V4-Pro-Max,themax- imum reasoning effort mode of DeepSeek-V4-Pro, establishes a new state-of-the-art among open-sourcelargelanguagemodels. AsdemonstratedbytheSimpleQA-Verified,DeepSeek-V4- Pro-Maxsignificantlyoutperformsallexistingopen-sourcebaselinesbyamarginof20absolute percentage points. Despite these advances, it currently trails the leading proprietary model, Gemini-3.1-Pro. Inthedomainofeducationalknowledgeandreasoning,DeepSeek-V4-Pro-Max marginallyoutperformsKimiandGLMacrosstheMMLU-Pro,GPQA,andHLEbenchmarks, althoughitlagsbehindleadingproprietarymodels. Broadly,DeepSeek-V4-Pro-Maxmarksa significantmilestoneinenhancingtheworldknowledgecapabilitiesofopen-sourcemodels. Inaddition,asignificantperformancegapexistsbetweenDeepSeek-V4-FlashandDeepSeek- V4-Pro on knowledge-based tasks; this is anticipated, as larger parameter counts facilitate greaterknowledgeretentionduringpre-training. Notably,bothmodelsdemonstrateimproved resultsonknowledgebenchmarkswhenallocatedhigherreasoningeffort. 38


Table7 | ComparisonamongdifferentsizesandmodesofDeepSeek-V4series. "Non-Think", "High",and"Max"denotereasoningeffort. DeepSeek-V4-Flash DeepSeek-V4-Pro Benchmark(Metric) Non-Think High Max Non-Think High Max gninosaeR&egdelwonK MMLU-Pro(EM) 83.0 86.4 86.2 82.9 87.1 87.5 SimpleQA-Verified(Pass@1) 23.1 28.9 34.1 45.0 46.2 57.9 Chinese-SimpleQA(Pass@1) 71.5 73.2 78.9 75.8 77.7 84.4 GPQADiamond(Pass@1) 71.2 87.4 88.1 72.9 89.1 90.1 HLE(Pass@1) 8.1 29.4 34.8 7.7 34.5 37.7 LiveCodeBench(Pass@1-COT) 55.2 88.4 91.6 56.8 89.8 93.5 Codeforces(Rating) - 2816 3052 - 2919 3206 HMMT2026Feb(Pass@1) 40.8 91.9 94.8 31.7 94.0 95.2 IMOAnswerBench(Pass@1) 41.9 85.1 88.4 35.3 88.0 89.8 Apex(Pass@1) 1.0 19.1 33.0 0.4 27.4 38.3 ApexShortlist(Pass@1) 9.3 72.1 85.7 9.2 85.5 90.2 gnoL MRCR1M(MMR) 37.5 76.9 78.7 44.7 83.3 83.5 CorpusQA1M(ACC) 15.5 59.3 60.5 35.6 56.5 62.0 citnegA TerminalBench2.0(Acc) 49.1 56.6 56.9 59.1 63.3 67.9 SWEVerified(Resolved) 73.7 78.6 79.0 73.6 79.4 80.6 SWEPro(Resolved) 49.1 52.3 52.6 52.1 54.4 55.4 SWEMultilingual(Resolved) 69.7 70.2 73.3 69.8 74.1 76.2 BrowseComp(Pass@1) - 53.5 73.2 - 80.4 83.4 HLEw/tools(Pass@1) - 40.3 45.1 - 44.7 48.2 MCPAtlasPublic(Pass@1) 64.0 67.4 69.0 69.4 74.2 73.6 GDPval-AA(Elo) - - 1395 - - 1554 Toolathlon(Pass@1) 40.7 43.5 47.8 46.3 49.0 51.8 Reasoning. DeepSeek-V4-Pro-Maxoutperformsallprioropenmodelsacrossreasoningbench- marks,andmatchesstate-of-the-artclosedmodelsonmanymetrics,whilethesmallerDeepSeek- V4-Flash-Max also surpasses the previous best open-source model, K2.6-Thinking, on code andmathreasoningtasks. Meanwhile,DeepSeek-V4-ProandDeepSeek-V4-Flashexcelincod- ing competitions. According to our evaluation, their performance is comparable to GPT-5.4, making this the first time an open model has matched a closed model on this task. On the Codeforcesleaderboard,DeepSeek-V4-Pro-Maxcurrentlyranks23rdamonghumancandidates. DeepSeek-V4alsodemonstratesstrongperformanceonformalmathematicaltaskunderboth agentic and compute-intensive settings. Under an agentic setup, it achieves state-of-the-art results,showninFigure8,outperformingpriormodelssuchasSeedProver(Chenetal.,2025). Withamorecompute-intensivepipeline,performancefurtherimproves,surpassingsystems includingAristotle(Achimetal.,2025)andmatchingthebestknownresultsunderthissetting. Agent. TheDeepSeek-V4seriesdemonstratesstrongagentperformanceinevaluations. For codeagenttasks,DeepSeek-V4-ProachievesresultscomparabletoK2.6andGLM-5.1,though all these open models still lag behind their closed-source counterparts. DeepSeek-V4-Flash underperformsDeepSeek-V4-Prooncodingtasks,particularlyonTerminalBench2.0. Asimilar trend is observed across other agent evaluations. It is worth noting that DeepSeek-V4-Pro performswellonMCPAtlasandToolathlon—twoevaluationtestsetsthatincludeawiderange oftoolsandMCPservices—indicatingthatourmodelhasexcellentgeneralizationcapability anddoesnotperformwellonlyoninternalframeworks. 39


PracticalRegime FrontierRegime Putnam-200Pass@8withminimaltools Putnam-2025withhybridformal-informal andboundedsampling. reasoningandsubstantialcomputescaling. Seed-1.5-Prover 26.50 Aristotle 100/120 Gemini-3-Pro 26.50 Seed-1.5-Prover 110/120 Seed-2.0-Pro 35.50 Axiom 120/120 DeepSeek-V4-Flash-Max 81.00 DeepSeek-V4 120/120 Figure 8 | Formal reasoning under practical and frontier regimes. Left: Putnam-200 Pass@8 evaluatesafixedrandomsubsetofPutnamBench(Tsoukalasetal.,2024)followingthesetup introducedbySeed-Prover;allmodelsaretestedonthesameproblemset. WefollowtheSeed- Proverprotocolbutreplaceproprietarysearchtoolswiththeopen-sourceLeanExplore(Asher, 2025),yieldingalightweightsettingwithminimalagenttoolsandboundedsampling. Right: Putnam-2025probesthefrontierofmathematicalreasoninginascaledhybridformal-informal regime, where informal reasoning is combined with formal verification to expose gaps and improverigor;DeepSeek-V4reachesaproof-perfect120/120. 1.0 0.8 0.6 0.4 0.2 0.0 8k 16k 32k 64k 128k 256k 512k 1024k Input Tokens RMM egarevA MRCR 8-needle 0.94 0.90 0.90 0.92 0.85 0.82 0.91 0.87 0.87 0.84 0.85 0.66 0.76 0.59 0.60 0.49 DeepSeek-V4-Pro-Max DeepSeek-V4-Flash-Max Figure9 | DeepSeek-V4seriesperformanceontheMRCRtask. 1M-TokenContext. DeepSeek-V4-ProoutperformsGemini-3.1-ProontheMRCRtask,which measuresin-contextretrieval,butremainsbehindClaudeOpus4.6. AsillustratedinFigure9, retrievalperformanceremainshighlystablewithina128Kcontextwindow. Whileaperformance degradation becomes visible beyond the 128K mark, the model’s retrieval capabilities at 1M tokensremainremarkablystrongcomparedtobothproprietaryandopen-sourcecounterparts. UnlikeMRCR,CorpusQAissimilartorealscenarios. Theevaluationresultsalsoindicatethat DeepSeek-V4-ProisbetterthanGemini-3.1-Pro. ReasoningEffort. AsshowninTable7,theMaxmode,whichemployslongercontextsand reduced length penalties in RL, outperforms the High mode on the most challenging tasks. Figure10presentsacomparisonofperformanceandcostamongDeepSeek-V4-Pro,DeepSeek- V4-Flash,andDeepSeek-V3.2onrepresentativereasoningandagentictasks. Byscalingtest-time compute,DeepSeek-V4seriesachievesubstantialimprovementsoverthepredecessor. Further- more,onreasoningtaskslikeHLE,DeepSeek-V4-Prodemonstrateshighertokenefficiencythan 40


40 35 30 25 20 15 10 5 0 20k 40k 60k 80k Total Tokens )%( 1@ssaP HLE Max 70 High Max Speciale 60 Think High 50 40 DeepSeek-V4-Pro DeepSeek-V4-Flash NNoonnee DeepSeek-V3.2 None 30 20k 30k 40k 50k Total Tokens )%( 1@ssaP TerminalBench 2.0 Max High None High Max Think None DeepSeek-V4-Pro DeepSeek-V4-Flash None DeepSeek-V3.2 Figure 10 | HLE and Terminal Bench 2.0 performance by reasoning effort. “None” indicates Non-thinkmode,and“Speciale”indicatesDeepSeek-V3.2-Specialemodel. DeepSeek-V3.2. 5.4. PerformanceonReal-WorldTasks Standardized benchmarks often struggle to capture the complexities of diverse, real-world tasks,creatingagapbetweentestresultsandactualuserexperience. Tobridgethis,wehave developedproprietaryinternalmetricsthatprioritizereal-worldusagepatternsovertraditional benchmarks. This approach ensures that our optimizations translate into tangible benefits. OurevaluationframeworkspecificallytargetstheprimaryusecasesoftheDeepSeekAPIand Chatbot,aligningmodelperformancewithpracticaldemands. 5.4.1. ChineseWriting One of the primary use cases for DeepSeek is Chinese writing. We conducted a rigorous evaluationonfunctionalwritingandcreativewriting. Table12presentsapairwisecomparison betweenDeepSeek-V4-ProandGemini-3.1-Proonfunctionalwritingtasks. Thesetasksconsist of common daily writing queries, where prompts are typically concise and straightforward. Gemini-3.1-Prowasselectedasthebaseline,asitstandsasthetop-performingexternalmodel forChinesewritinginourevaluations. TheresultsindicatethatDeepSeek-V4-Prooutperforms thebaselinewithanoverallwinrateof62.7%versus34.1%;thisisprimarilybecauseGemini occasionallyallowsitsinherentstylisticpreferencestooverridetheuser’sexplicitrequirements inChinesewritingscenarios. Table 13 presents the creative writing comparison, which is evaluated along two axes: instructionfollowingandwritingquality. ComparedwithGemini-3.1-Pro,DeepSeek-V4-Pro achievesa60.0%winrateininstructionfollowingand77.5%inwritingquality,demonstrating a marginal improvement in instruction following and a substantial gain in writing quality. AlthoughDeepSeek-V4-Proyieldssuperiorresultsinaggregateusercaseanalysis,anevaluation restricted to the most challenging prompts — specifically those involving high-complexity constraints or multi-turn scenarios — reveals that Claude Opus 4.5 retains a performance advantageoverDeepSeek-V4-Pro. AsshowninTable14,ClaudeOpus4.5achievesa52.0%win rateversus45.9%. 41


5.4.2. Search Search-augmented question answering is a core capability of the DeepSeek chatbot. On the DeepSeekwebandapp,the"non-think"modeemploysRetrieval-AugmentedSearch(RAG), whereasthe"thinking"modeutilizesagenticsearch. RetrievalAugmentedSearch. WeconductedapairwiseevaluationcomparingDeepSeek-V4- ProandDeepSeek-V3.2acrossbothobjectiveandsubjectiveQ&Acategories. Aspresentedin Table11,DeepSeek-V4-ProoutperformsDeepSeek-V3.2byasubstantialmargin,demonstrating a consistent advantage across both categories. The most pronounced gains are observed in single-value search and planning & strategy tasks, suggesting that DeepSeek-V4-Pro excels atlocating precise factualanswers and synthesizingstructuredplans from retrievedcontext. However,DeepSeek-V3.2remainsrelativelycompetitiveoncomparisonandrecommendation tasks,indicatingpotentialroomforimprovementforDeepSeek-V4-Proinscenariosrequiring balanced,multi-perspectivereasoningoversearchresults. Agentic Search. Unlike standard RAG, agentic search empowers the model to iteratively invokesearchandfetchtoolsperquery,significantlyenhancingoverallsearchperformance. For thethinkingmodeinDeepSeek-Chat,weoptimizedtheagenticsearchfunctiontomaximize responseaccuracywithinapredefined"thinkingbudget". AsshowninTable9,agenticsearch consistentlyoutperformsRAG,particularlyoncomplextasks. Furthermore, itscostremains highlyefficient,withagenticsearchbeingonlymarginallymoreexpensivethanstandardRAG (seeTable10). 5.4.3. White-CollarTask Torigorouslyevaluatethemodel’sutilityinsophisticatedenterpriseproductivityscenarios,we constructedacomprehensivesuiteof30advancedChineseprofessionaltasks. Theseworkflows deliberatelyencompasshigh-levelcognitivedemands,includingin-depthinformationanalysis, comprehensive document generation, and nuanced document editing, spanning a diverse spectrumof13criticalindustries(e.g.,finance,education,law,andtechnology). Theevaluation wasconductedwithinanin-houseagentharnessequippedwithbasictools,includingBashand websearch. Giventheopen-endednatureofthesetasks,automatedmetricsusuallyfallshortincapturing thenuancesofahigh-qualityresponse. Therefore,weconductedhumanevaluationstocompare theperformanceofDeepSeek-V4-Pro-MaxagainstOpus-4.6-Max. Annotatorsblindlyassessed themodeloutputsacrossfourdimensions: • TaskCompletion: Whetherthecoreproblemwassuccessfullyresolved. • InstructionFollowing: Adherencetospecificconstraintsanddirectives. • ContentQuality: Factualaccuracy,logicalcoherence,andprofessionaltone. • FormattingAesthetics: Layoutreadabilityandvisualpresentation. AsillustratedinFigure11,DeepSeek-V4-Pro-MaxoutperformsOpus-4.6-Maxondiverse Chinesewhite-collartasks,achievinganimpressivenon-lossrateof63%,anddemonstrating consistentadvantagesacrossanalysis,generation,andeditingtasks. Thedetaileddimension scores shown in Figure 12 highlight the model’s primary strengths in Task Completion and 42


Content Quality. Specifically, DeepSeek-V4-Pro-Max proactively anticipates implicit user in- tentsbyfrequentlyprovidingsupplementaryinsightsandself-verificationsteps. Italsoexcels in long-form generation, delivering in-depth, coherent narratives rather than relying on the overlysimplisticbulletpointsfrequentlyproducedbyOpus-4.6-Max. Additionally,themodel strictlyconformstoformalprofessionalconventions,suchasstandardizedChinesehierarchical numbering. However, in terms of Instruction Following, it occasionally overlooks specific formatting constraints and slightly trails Opus. Furthermore, the model is less proficient at condensing extensive text inputs into succinct summaries. Finally, its Formatting Aesthetics stillhavesubstantialroomforimprovementregardingtheoverallvisualdesignofpresentation slides. Figure 13, 14, and 15 present several test cases; due to the extensive length of certain outputs,onlypartialpagesaredisplayed. Win Rate: DeepSeek-V4-Pro-Max vs Opus-4.6-Max analysis 55.0% 8.0% 37.0% 100 95 generation 52.0% 10.0% 38.0% 90 85 editing 47.0% 18.0% 35.0% 80 overall 53.0% 10.0% 37.0% 75 70 0% 20% 40% Proportion 60% 80% 100% Task Completio In n struction Following Content Qua F li o ty rmatting Aesthetics Overall Win Tie Lose Figure11 | Win-ratecomparisonacrossanaly- sis,generation,editingtasks,andtheoverall performance. erocS Score: DeepSeek-V4-Pro-Max vs Opus-4.6-Max 98.32 96.68 88.88 87.76 86.52 83.32 84.06 78.00 76.68 72.68 DeepSeek-V4-Pro-Max Opus-4.6-Max Figure12 | Detaileddimensionscoresinclud- ingTaskCompletion,ContentQuality,Format- tingAesthetics,andInstructionFollowing. Figure13 | Exampleoutputofataskwhichrequiresdraftingajointmarketingproposalfora popularbubbleteabrandandtheBeijingSubway. 43


5.4.4. CodeAgent Tobenchmarkourcodingagentcapability,wecuratetasksfromrealinternalR&Dworkloads Wecollect∼200challengingtasksfrom50+internalengineers,spanningfeaturedevelopment, bug fixing, refactoring, and diagnostics across diverse technology stacks including PyTorch, CUDA,Rust,andC++. Eachtaskisaccompaniedbyitsoriginalrepository,thecorresponding executionenvironment,andhuman-annotatedscoringrubrics;afterrigorousqualityfiltering, 30tasksareretainedastheevaluationset. AsshowninTable8,DeepSeek-V4-Prosignificantly outperformsClaudeSonnet4.5andapproachesthelevelofClaudeOpus4.5. Table8 | ComparisononR&DCodingBenchmark(externalmodelsincludedstrictlyforevalua- tionpurposes). Opus4.5 Opus4.6 Model Haiku4.5 Sonnet4.5 DeepSeek-V4-Pro-Max Opus4.5 Thinking Thinking PassRate(%) 13 47 67 70 73 80 InasurveyaskingDeepSeekdevelopersandresearchers(𝑁 =85)—allwithexperienceof usingDeepSeek-V4-Proforagenticcodingintheirdailywork—whetherDeepSeek-V4-Prois readytoserveastheirdefaultandprimarycodingmodelcomparedtootherfrontiermodels,52% saidyes,39%leanedtowardyes,andfewerthan9%saidno. RespondentsfindDeepSeek-V4-Pro todeliversatisfactoryresultsacrossmosttasks,butnotetrivialmistakes,misinterpretationof vagueprompts,andoccasionalover-thinking. 6. Conclusion, Limitations, and Future Directions Inthiswork,wepresentapreviewversionofDeepSeek-V4series,aimingatnext-generation largelanguagemodelsthatbreaktheefficiencybarrierofultra-long-contextprocessing. Bycom- biningahybridattentionarchitecturethatintegratesCSAandHCA,DeepSeek-V4seriesachieve adramaticleapinlong-sequenceefficiency. Thearchitecturalinnovations,togetherwithexten- siveinfrastructureoptimization,enableefficientnativesupportformillion-tokencontextsand establishanecessaryfoundationforfuturetest-timescaling,long-horizontasks,andemerging paradigmssuchasonlinelearning. EvaluationresultsdemonstratethatDeepSeek-V4-Pro-Max, themaximumreasoningeffortmodeofDeepSeek-V4-Pro,redefinesthestate-of-the-artforopen models. It substantially outperforms prior open-source models on knowledge benchmarks, achievessuperiorreasoningperformanceclosetothefrontierproprietarymodels,anddelivers competitiveagentcapabilities. Meanwhile,DeepSeek-V4-Flash-Maxattainscomparablereason- ingperformancetoleadingclosedmodelswhilemaintainingahighlycost-efficientarchitecture. WebelieveDeepSeek-V4seriesusherinaneweraofmillion-lengthcontextsforopenmodels andpavethewaytowardbetterefficiency,scale,andintelligence. Inpursuitofextremelong-contextefficiency,DeepSeek-V4seriesadoptedaboldarchitec- tural design. To minimize risk, we retained many preliminarily validated components and tricks, which, while effective, made the architecture relatively complex. In future iterations, wewillcarryoutmorecomprehensiveandprincipledinvestigationstodistillthearchitecture down to its most essential designs, making it more elegant without sacrificing performance. Meanwhile,althoughAnticipatoryRoutingandSwiGLUClampinghavebeenproveneffective inmitigatingtraininginstabilities,theirunderlyingprinciplesremaininsufficientlyunderstood. Wewillactivelystudyfoundationalproblemsontrainingstabilityandstrengtheninternalmetric monitoring,aimingforamoreprincipledandpredictiveapproachtostablelarge-scaletraining. 44


Inaddition,beyondtheMoEandsparseattentionarchitecture,wewillalsoproactivelyexplore model sparsity along new dimensions — such as more sparse embedding modules (Cheng etal.,2026)—tofurtherimprovecomputationalandmemoryefficiencywithoutcompromising capability. We will also continuously investigate low-latency architectures and system tech- niquestomakelong-contextdeploymentandinteractionmoreresponsive. Furthermore,we recognizetheimportanceandpracticalvalueoflong-horizon,multi-roundagentictasks,and will continue to iterate and explore in this direction. We are also working on incorporating multimodal capabilities to our models. Finally, we are committed to developing better data curationandsynthesisstrategiestoconsistentlyenhancemodelintelligence,robustness,and practicalusabilityacrossanincreasinglybroadrangeofscenariosandtasks. References AA. Gdpval-aaleaderboard,2025. URLhttps://artificialanalysis.ai/methodolog y/intelligence-benchmarking#gdpval-aa. T.Achim,A.Best,A.Bietti,K.Der,M.Fédérico,S.Gukov,D.Halpern-Leistner,K.Henningsgard, Y. Kudryashov, A. Meiburg, et al. Aristotle: Imo-level automated theorem proving. arXiv preprintarXiv:2510.01346,2025. A.Agache,M.Brooker,A.Florescu,A.Iordache,A.Liguori,R.Neugebauer,P.Piwonka,and D.-M.Popa. Firecracker: lightweightvirtualizationforserverlessapplications. InProceedings ofthe17thUsenixConferenceonNetworkedSystemsDesignandImplementation,NSDI’20, page419–434,USA,2020.USENIXAssociation. ISBN9781939133137. O.J.Aimuyo,B.Oh,andR.Singh. Flashmoe: Fastdistributedmoeinasinglekernel. Advances inNeuralInformationProcessingSystems,2025. URLhttps://neurips.cc/virtual/2 025/poster/119124. J.Ainslie,J.Lee-Thorp,M.deJong,Y.Zemlyanskiy,F.Lebrón,andS.Sanghai. Gqa: Training generalizedmulti-querytransformermodelsfrommulti-headcheckpoints. arXivpreprint arXiv:2305.13245,2023. J.Asher. LeanExplore: AsearchengineforLean4declarations,2025. URLhttps://arxiv.or g/abs/2506.11085. Y.Bai,Y.Bao,G.Chen,J.Chen,N.Chen,R.Chen,Y.Chen,Y.Chen,Y.Chen,Z.Chen,J.Cui, H.Ding,M.Dong,A.Du,C.Du,D.Du,Y.Du,Y.Fan,Y.Feng,K.Fu,B.Gao,H.Gao,P.Gao, T.Gao,X.Gu,L.Guan,H.Guo,J.Guo,H.Hu,X.Hao,T.He,W.He,W.He,C.Hong,Y.Hu, Z.Hu,W.Huang,Z.Huang,Z.Huang,T.Jiang,Z.Jiang,X.Jin,Y.Kang,G.Lai,C.Li,F.Li, H.Li,M.Li,W.Li,Y.Li,Y.Li,Z.Li,Z.Li,H.Lin,X.Lin,Z.Lin,C.Liu,C.Liu,H.Liu,J.Liu, J.Liu,L.Liu,S.Liu,T.Y.Liu,T.Liu,W.Liu,Y.Liu,Y.Liu,Y.Liu,Y.Liu,Z.Liu,E.Lu,L.Lu, S.Ma,X.Ma,Y.Ma,S.Mao,J.Mei,X.Men,Y.Miao,S.Pan,Y.Peng,R.Qin,B.Qu,Z.Shang, L.Shi,S.Shi,F.Song,J.Su,Z.Su,X.Sun,F.Sung,H.Tang,J.Tao,Q.Teng,C.Wang,D.Wang, F. Wang, and H. Wang. Kimi K2: open agentic intelligence. CoRR, abs/2507.20534, 2025a. URLhttps://doi.org/10.48550/arXiv.2507.20534. Y.Bai,S.Tu,J.Zhang,H.Peng,X.Wang,X.Lv,S.Cao,J.Xu,L.Hou,Y.Dong,etal. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume1: LongPapers),pages3639–3664,2025b. 45


M.Balunovic´,J.Dekoninck,I.Petrov,N.Jovanovic´,andM.Vechev. Matharena: Evaluatingllms on uncontaminated math competitions. Proceedings of the Neural Information Processing SystemsTrackonDatasetsandBenchmark,2025. C.Bandi,B.Hertzberg,G.Boo,T.Polakam,J.Da,S.Hassaan,M.Sharma,A.Park,E.Hernandez, D. Rambado, et al. Mcp-atlas: A large-scale benchmark for tool-use competency with real mcpservers. arXivpreprintarXiv:2602.00933,2026. F. Bellard. Qemu, a fast and portable dynamic translator. In Proceedings of the Annual ConferenceonUSENIXAnnualTechnicalConference,ATEC’05,page41,USA,2005.USENIX Association. I.Bello,H.Pham,Q.V.Le,M.Norouzi,andS.Bengio. Neuralcombinatorialoptimizationwith reinforcementlearning,2017. URLhttps://openreview.net/forum?id=rJY3vK9eg. J. Chen, W.Chen, J. Du, J. Hu, Z. Jiang, A. Jie, X. Jin, X.Jin, C. Li, W. Shi, Z. Wang, M.Wang, C. Wei, S. Wei, H. Xin, F. Yang, W. Gao, Z. Yuan, T. Zhan, Z. Zheng, T. Zhou, and T. H. Zhu. Seed-prover 1.5: Mastering undergraduate-level theorem proving via learning from experience,2025. URLhttps://arxiv.org/abs/2512.17260. M.Chen,J.Tworek,H.Jun,Q.Yuan,H.P.deOliveiraPinto,J.Kaplan,H.Edwards,Y.Burda, N.Joseph,G.Brockman,A.Ray,R.Puri,G.Krueger,M.Petrov,H.Khlaaf,G.Sastry,P.Mishkin, B.Chan,S.Gray,N.Ryder,M.Pavlov,A.Power,L.Kaiser,M.Bavarian,C.Winter,P.Tillet, F.P.Such,D.Cummings,M.Plappert,F.Chantzis,E.Barnes,A.Herbert-Voss,W.H.Guss, A.Nichol,A.Paino,N.Tezak,J.Tang,I.Babuschkin,S.Balaji,S.Jain,W.Saunders,C.Hesse, A.N.Carr,J.Leike,J.Achiam,V.Misra,E.Morikawa,A.Radford,M.Knight,M.Brundage, M.Murati,K.Mayer,P.Welinder,B.McGrew,D.Amodei,S.McCandlish,I.Sutskever,and W.Zaremba. Evaluatinglargelanguagemodelstrainedoncode. CoRR,abs/2107.03374,2021. URLhttps://arxiv.org/abs/2107.03374. T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy. TVM: An automated End-to-End optimizing com- piler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pages 578–594, Carlsbad, CA, Oct. 2018. USENIX Association. ISBN978-1-939133-08-3. URLhttps://www.usenix.org/conference/osdi18/prese ntation/chen. A.Cheng,A.Jacovi,A.Globerson,B.Golan,C.Kwong,C.Alberti,C.Tao,E.Ben-David,G.S. Tomar,L.Haas,etal. Thefactsleaderboard: Acomprehensivebenchmarkforlargelanguage modelfactuality. arXivpreprintarXiv:2512.10791,2025. X.Cheng,W.Zeng,D.Dai,Q.Chen,B.Wang,Z.Xie,K.Huang,X.Yu,Z.Hao,Y.Li,H.Zhang, H.Zhang,D.Zhao,andW.Liang. Conditionalmemoryviascalablelookup: Anewaxisof sparsityforlargelanguagemodels. CoRR,abs/2601.07372,2026. doi: 10.48550/ARXIV.2601. 07372. URLhttps://doi.org/10.48550/arXiv.2601.07372. K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J.Hilton, R.Nakano, etal. Trainingverifierstosolvemathwordproblems. arXivpreprint arXiv:2110.14168,2021. D.Dai,C.Deng,C.Zhao,R.X.Xu,H.Gao,D.Chen,J.Li,W.Zeng,X.Yu,Y.Wu,Z.Xie,Y.K. Li,P.Huang,F.Luo,C.Ruan,Z.Sui,andW.Liang. Deepseekmoe: Towardsultimateexpert specialization in mixture-of-experts language models. CoRR, abs/2401.06066, 2024. URL https://doi.org/10.48550/arXiv.2401.06066. 46


T.Dao,D.Haziza,F.Massa,andG.Sizov. Flash-decodingforlong-contextinference,2023. URL https://pytorch.org/blog/flash-decoding/. L. De Moura and N. Bjørner. Z3: an efficient smt solver. In Proceedings of the Theory and Practice of Software, 14th International Conference on Tools and Algorithms for the ConstructionandAnalysisofSystems,TACAS’08/ETAPS’08,page337–340,Berlin,Heidel- berg,2008.Springer-Verlag. ISBN3540787992. DeepSeek-AI. Deepseek-coder-v2: Breakingthebarrierofclosed-sourcemodelsincodeintelli- gence. CoRR,abs/2406.11931,2024. URLhttps://doi.org/10.48550/arXiv.2406.11 931. DeepSeek-AI. Deepseek-v3technicalreport. CoRR,abs/2412.19437,2024. URLhttps://doi. org/10.48550/arXiv.2412.19437. DeepSeek-AI. Deepseek-v2: Astrong,economical,andefficientmixture-of-expertslanguage model. CoRR,abs/2405.04434,2024. URLhttps://doi.org/10.48550/arXiv.2405.04 434. DeepSeek-AI. Fire-flyerfilesystem,2025. URLhttps://github.com/deepseek-ai/3FS. DeepSeek-AI. Deepseek-r1incentivizesreasoninginllmsthroughreinforcementlearning. Nat., 645(8081):633–638,2025. URLhttps://doi.org/10.1038/s41586-025-09422-z. DeepSeek-AI. Deepseek-v3.2: Pushingthefrontierofopenlargelanguagemodels,2025. URL https://arxiv.org/abs/2512.02556. X.Deng,J.Da,E.Pan,Y.Y.He,C.Ide,K.Garg,N.Lauffer,A.Park,N.Pasari,C.Rane,K.Sampath, M.Krishnan, S.Kundurthy, S.Hendryx,Z.Wang,V.Bharadwaj,J.Holm,R.Aluri,C.B.C. Zhang,N.Jacobson,B.Liu,andB.Kenstler. Swe-benchpro: Canaiagentssolvelong-horizon softwareengineeringtasks?,2025. URLhttps://arxiv.org/abs/2509.16941. H.Ding,Z.Wang,G.Paolini,V.Kumar,A.Deoras,D.Roth,andS.Soatto. Fewertruncations improvelanguagemodeling. arXivpreprintarXiv:2404.10830,2024. X.Dong,Y.Fu,S.Diao,W.Byeon,Z.CHEN,A.S.Mahabaleshwarkar,S.-Y.Liu,M.V.keirsbilck, M.-H.Chen,Y.Suhara,Y.C.Lin,J.Kautz,andP.Molchanov. Hymba: Ahybrid-headarchi- tectureforsmalllanguagemodels. InTheThirteenthInternationalConferenceonLearning Representations,2025. URLhttps://openreview.net/forum?id=A1ztozypga. X.Du,Y.Yao,K.Ma,B.Wang,T.Zheng,K.Zhu,M.Liu,Y.Liang,X.Jin,Z.Wei,etal. Supergpqa: Scalingllmevaluationacross285graduatedisciplines. arXivpreprintarXiv:2502.14739,2025. D.Dua,Y.Wang,P.Dasigi,G.Stanovsky,S.Singh,andM.Gardner. DROP:Areadingcompre- hensionbenchmarkrequiringdiscretereasoningoverparagraphs.InJ.Burstein,C.Doran,and T.Solorio,editors,Proceedingsofthe2019ConferenceoftheNorthAmericanChapterofthe Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019,Minneapolis,MN,USA,June2-7,2019,Volume1(LongandShortPapers),pages2368– 2378.AssociationforComputationalLinguistics, 2019. doi: 10.18653/V1/N19-1246. URL https://doi.org/10.18653/v1/n19-1246. X.Gao,M.Dong,X.Miao,W.Du,C.Yu,andH.Chen. Erofs: acompression-friendlyreadonly file system for resource-scarce devices. In Proceedings of the 2019 USENIX Conference on UsenixAnnualTechnicalConference,USENIXATC’19,page149–162,USA,2019.USENIX Association. ISBN9781939133038. 47


A.P.Gema, J.O.J.Leang, G.Hong, A.Devoto, A.C.M.Mancino, R.Saxena, X.He, Y.Zhao, X. Du, M. R. G. Madani, C. Barale, R. McHardy, J. Harris, J. Kaddour, E. van Krieken, and P.Minervini. Arewedonewithmmlu? CoRR,abs/2406.04127,2024. URLhttps://doi.or g/10.48550/arXiv.2406.04127. F. Gloeckle, B. Y. Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve. Better & faster large language models via multi-token prediction. In Forty-first International Conference on MachineLearning,ICML2024,Vienna,Austria,July21-27,2024.OpenReview.net,2024. URL https://openreview.net/forum?id=pEWAcejiU2. L. Haas, G. Yona, G. D’Antonio, S. Goldshtein, and D. Das. Simpleqa verified: A reliable factuality benchmark to measure parametric knowledge. arXiv preprint arXiv:2509.07968, 2025. Y. He, S. Li, J. Liu, Y. Tan, W. Wang, H. Huang, X. Bu, H. Guo, C. Hu, B. Zheng, et al. Chi- nese simpleqa: A chinese factuality evaluation for large language models. arXiv preprint arXiv:2411.07140,2024. D.Hendrycks,C.Burns,S.Basart,A.Zou,M.Mazeika,D.Song,andJ.Steinhardt. Measuring massivemultitasklanguageunderstanding. arXivpreprintarXiv:2009.03300,2020. D.Hendrycks,C.Burns,S.Kadavath,A.Arora,S.Basart,E.Tang,D.Song,andJ.Steinhardt.Mea- suringmathematicalproblemsolvingwiththemathdataset. arXivpreprintarXiv:2103.03874, 2021. Y.Huang,Y.Bai,Z.Zhu,J.Zhang,J.Zhang,T.Su,J.Liu,C.Lv,Y.Zhang,J.Lei,etal. C-Eval: A multi-levelmulti-disciplinechineseevaluationsuiteforfoundationmodels. arXivpreprint arXiv:2305.08322,2023. D.HupkesandN.Bogoychev. Multiloko: amultilinguallocalknowledgebenchmarkforllms spanning31languages. CoRR,abs/2504.10356,2025. doi: 10.48550/ARXIV.2504.10356. URL https://doi.org/10.48550/arXiv.2504.10356. B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko. Quantizationandtrainingofneuralnetworksforefficientinteger-arithmetic-onlyinference. InProceedingsoftheIEEEConferenceonComputerVisionandPatternRecognition(CVPR), June2018. N.Jain,K.Han,A.Gu,W.-D.Li,F.Yan,T.Zhang,S.Wang,A.Solar-Lezama,K.Sen,andI.Stoica. Livecodebench: Holisticandcontaminationfreeevaluationoflargelanguagemodelsforcode. arXivpreprintarXiv:2403.07974,2024. K.Jordan,Y.Jin,V.Boza,J.You,F.Cesista,L.Newhouse,andJ.Bernstein. Muon: Anoptimizer forhiddenlayersinneuralnetworks. Citedon,page10,2024. M.Joshi,E.Choi,D.Weld,andL.Zettlemoyer. TriviaQA:Alargescaledistantlysupervisedchal- lengedatasetforreadingcomprehension.InR.BarzilayandM.-Y.Kan,editors,Proceedingsof the55thAnnualMeetingoftheAssociationforComputationalLinguistics(Volume1: Long Papers), pages 1601–1611, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1147. URLhttps://aclanthology.org/P17-1147. H. Li, Y. Yuan, R. Du, K. Ma, L. Liu, and W. Hsu. DADI: Block-Level image service for agile andelasticapplicationdeployment. In2020USENIXAnnualTechnicalConference(USENIX ATC 20), pages 727–740. USENIX Association, July 2020. ISBN 978-1-939133-14-4. URL https://www.usenix.org/conference/atc20/presentation/li-huiba. 48


H.Li,Y.Zhang,F.Koto,Y.Yang,H.Zhao,Y.Gong,N.Duan,andT.Baldwin. CMMLU:Measur- ingmassivemultitasklanguageunderstandinginChinese. arXivpreprintarXiv:2306.09212, 2023. J.Li, W.Zhao, J.Zhao, W.Zeng, H.Wu, X.Wang, R.Ge, Y.Cao, Y.Huang, W.Liu, etal. The tooldecathlon: Benchmarkinglanguageagentsfordiverse,realistic,andlong-horizontask execution. arXivpreprintarXiv:2510.25726,2025. Y. Li, F. Wei, C. Zhang, and H. Zhang. EAGLE: speculative sampling requires rethinking featureuncertainty. InForty-firstInternationalConferenceonMachineLearning,ICML2024, Vienna,Austria,July21-27,2024.OpenReview.net,2024. URLhttps://openreview.net /forum?id=1NdN7eXyb4. J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y. Du, Y. Qin, W. Xu, E. Lu, J. Yan, Y. Chen, H. Zheng, Y.Liu,S.Liu,B.Yin,W.He,H.Zhu,Y.Wang,J.Wang,M.Dong,Z.Zhang,Y.Kang,H.Zhang, X. Xu, Y. Zhang, Y. Wu, X. Zhou, and Z. Yang. Muon is scalable for LLM training. CoRR, abs/2502.16982,2025. URLhttps://doi.org/10.48550/arXiv.2502.16982. I. Loshchilov and F. Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101,2017. K.LuandT.M.Lab. On-policydistillation. ThinkingMachinesLab: Connectionism,2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/on-policy-distillation. Z.Lu,C.Li,Y.Shi,W.Shen,M.Yan,andF.Huang. Corpusqa: A10milliontokenbenchmark forcorpus-levelanalysisandreasoning. arXivpreprintarXiv:2601.14952,2026. T. Luong, D. Hwang, H. H. Nguyen, G. Ghiasi, Y. Chervonyi, I. Seo, J. Kim, G. Bingham, J. Lee, S. Mishra, A. Zhai, C. H. Hu, H. Michalewski, J. Kim, J. Ahn, J. Bae, X. Song, T. H. Trinh, Q. V. Le, and J. Jung. Towards robust mathematical reasoning. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025. URL https://aclanthology.org/2025.emnlp-main.1794/. M.A.Merrill,A.G.Shaw,N.Carlini,B.Li,H.Raj,I.Bercovich,L.Shi,J.Y.Shin,T.Walshe,E.K. Buchanan,etal. Terminal-bench: Benchmarkingagentsonhard,realistictasksincommand lineinterfaces. arXivpreprintarXiv:2601.11868,2026. MiniMax. Meetminimax-m2,2025. URLhttps://github.com/MiniMax-AI/MiniMax-M2. L. d. Moura and S. Ullrich. The lean 4 theorem prover and programming language. In InternationalConferenceonAutomatedDeduction,pages625–635.Springer,2021. Y.Nesterov. Amethodofsolvingaconvexprogrammingproblemwithconvergencerate𝑂(1/𝑘2). SovietMathematicsDoklady,27:372–376,1983. NVIDIACorporation. cublasdocumentation,2024. URLhttps://docs.nvidia.com/cuda /cublas/. Version12.4.Accessed: 2024-09-16. OpenAI. Multilingual massive multitask language understanding (mmmlu), 2024a. URL https://huggingface.co/datasets/openai/MMMLU. OpenAI. Openai mrcr: Long context multiple needle in a haystack benchmark, 2024b. URL https://huggingface.co/datasets/openai/mrcr. 49


OpenAI. Learningtoreasonwithllms,2024c. URLhttps://openai.com/index/learnin g-to-reason-with-llms. OpenAI. IntroducingSimpleQA,2024d. URLhttps://openai.com/index/introducing -simpleqa/. OpenAI. Introducing SWE-bench verified we’re releasing a human-validated subset of swe- benchthatmore,2024e. URLhttps://openai.com/index/introducing-swe-bench-v erified/. OpenAI. gpt-oss-120b&gpt-oss-20bmodelcard. CoRR,abs/2508.10925,2025. doi: 10.48550/A RXIV.2508.10925. URLhttps://doi.org/10.48550/arXiv.2508.10925. M.Osama,D.Merrill,C.Cecka,M.Garland,andJ.D.Owens. Stream-k: Work-centricparallel decompositionfordensematrix-matrixmultiplicationonthegpu. InProceedingsofthe28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, pages429–431,2023. T. Patwardhan, R. Dias, E. Proehl, G. Kim, M. Wang, O. Watkins, S. P. Fishman, M. Aljubeh, P. Thacker, L. Fauconnet, et al. Gdpval: Evaluating ai model performance on real-world economicallyvaluabletasks. arXivpreprintarXiv:2510.04374,2025. L.Phan,A.Gatti,Z.Han,N.Li,J.Hu,H.Zhang,C.B.C.Zhang,M.Shaaban,J.Ling,S.Shi,etal. Humanity’slastexam. arXivpreprintarXiv:2501.14249,2025. W.Qi,Y.Yan,Y.Gong,D.Liu,N.Duan,J.Chen,R.Zhang,andM.Zhou. Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training. In T. Cohn, Y. He, and Y. Liu, edi- tors,FindingsoftheAssociationforComputationalLinguistics: EMNLP2020,OnlineEvent, 16-20November2020,volumeEMNLP2020ofFindingsofACL,pages2401–2410.Associa- tionforComputationalLinguistics,2020. URLhttps://doi.org/10.18653/v1/2020.f indings-emnlp.217. Qwen. Qwen3technicalreport. CoRR,abs/2505.09388,2025. doi: 10.48550/ARXIV.2505.09388. URLhttps://doi.org/10.48550/arXiv.2505.09388. S.Rajbhandari,J.Rasley,O.Ruwase,andY.He.Zero: Memoryoptimizationstowardtrainingtril- lionparametermodels. InSC20: InternationalConferenceforHighPerformanceComputing, Networking,StorageandAnalysis,pages1–16.IEEE,2020. J.K.Reed,Z.DeVito,H.He,A.Ussery,andJ.Ansel. Torch.fx: Practicalprogramcaptureand transformationfordeeplearninginpython,2022. URLhttps://arxiv.org/abs/2112.0 8429. D.Rein,B.L.Hou,A.C.Stickland,J.Petty,R.Y.Pang,J.Dirani,J.Michael,andS.R.Bowman. GPQA:Agraduate-levelgoogle-proofq&abenchmark. arXivpreprintarXiv:2311.12022,2023. G. T. M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B.Shahriari,A.Ram’e,J.Ferret,P.Liu,P.D.Tafti,A.Friesen,M.Casbon,S.Ramos,R.Kumar, C.L.Lan,S.Jerome,A.Tsitsulin,N.Vieillard,P.Stan´czyk,S.Girgin,N.Momchev,M.Hoff- man, S. Thakoor, J.-B. Grill, B. Neyshabur, A. Walton, A. Severyn, A. Parrish, A. Ahmad, A. Hutchison, A. Abdagic, A. Carl, A. Shen, A. Brock, A. Coenen, A. Laforge, A. Paterson, B.Bastian,B.Piot,B.Wu,B.Royal,C.Chen,C.Kumar,C.Perry,C.A.Welty,C.A.Choquette- Choo,D.Sinopalnikov,D.Weinberger,D.Vijaykumar,D.Rogozi’nska,D.Herbison,E.Bandy, E.Wang,E.Noland,E.Moreira,E.Senter,E.Eltyshev,F.Visin,G.Rasskin,G.Wei,G.Cameron, 50


G. Martins, H. Hashemi, H. Klimczak-Pluci’nska, H. Batra, H. Dhand, I. Nardini, J. Mein, J.Zhou,J.Svensson,J.Stanway,J.Chan,J.Zhou,J.Carrasqueira,J.Iljazi,J.Becker,J.Fernan- dez,J.R.vanAmersfoort,J.Gordon,J.Lipschultz,J.Newlan,J.Ji,K.Mohamed,K.Badola, K.Black,K.Millican,K.McDonell,K.Nguyen,K.Sodhia,K.Greene,L.L.Sjoesund,L.Usui, L.Sifre,L.Heuermann,L.ciaLago,L.McNealus,L.B.Soares,L.Kilpatrick,L.Dixon,L.L.B. Martins, M. Reid, M. Singh, M. Iverson, M. Gorner, M. Velloso, M. Wirth, M. Davidow, M.Miller,M.Rahtz,M.Watson,M.Risdal,M.Kazemi,M.Moynihan,M.Zhang,M.Kahng, M.Park,M.Rahman,M.Khatwani,N.Dao,N.shadBardoliwalla,N.Devanathan,N.Dumai, N.Chauhan,O.Wahltinez,P.Botarda,P.Barnes,P.Barham,P.Michel,P.chongJin,P.Georgiev, P.Culliton,P.Kuppala,R.Comanescu,R.Merhej,R.Jana,R.A.Rokni,R.Agarwal,R.Mullins, S. Saadat, S. M. M. Carthy, S. Perrin, S. M. R. Arnold, S. bastian Krause, S. Dai, S. Garg, S.Sheth,S.Ronstrom,S.Chan,T.Jordan,T.Yu,T.Eccles,T.Hennigan,T.Kociský,T.Doshi, V.Jain,V.Yadav,V.Meshram,V.Dharmadhikari,W.Barkley,W.Wei,W.Ye,W.Han,W.Kwon, X.Xu,Z.Shen,Z.Gong,Z.Wei,V.Cotruta,P.Kirk,A.Rao,M.Giang,L.Peran,T.Warkentin, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, D. Sculley, J. Banks, A. Dragan, S. Petrov, O.Vinyals,J.Dean,D.Hassabis,K.Kavukcuoglu,C.Farabet,E.Buchatskaya,S.Borgeaud, N. Fiedel, A. Joulin, K. Kenealy, R. Dadashi, and A. Andreev. Gemma 2: Improving open languagemodelsatapracticalsize. arXivpreprintarXiv:2408.00118,2024. S. Roller, S. Sukhbaatar, A. Szlam, and J. Weston. Hash layers for large sparse models. In M.Ranzato,A.Beygelzimer,Y.N.Dauphin,P.Liang,andJ.W.Vaughan,editors,Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 17555–17566, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/92bf5e6240737 e0326ea59846a83e076-Abstract.html. B.D.Rouhani,R.Zhao,A.More,M.Hall,A.Khodamoradi,S.Deng,D.Choudhary,M.Cornea, E.Dellinger,K.Denolf,S.Dusan,V.Elango,M.Golub,A.Heinecke,P.James-Roxby,D.Jani, G.Kolhe,M.Langhammer,A.Li,L.Melnick,M.Mesmakhosroshahi,A.Rodriguez,M.Schulte, R. Shafipour, L. Shao, M. Siu, P. Dubey, P. Micikevicius, M. Naumov, C. Verrilli, R. Wittig, D.Burger,andE.Chung. Microscalingdataformatsfordeeplearning,2023. K.Sakaguchi,R.L.Bras,C.Bhagavatula,andY.Choi. Winogrande: Anadversarialwinograd schemachallengeatscale,2019. Z.Shao,Y.Luo,C.Lu,Z.Z.Ren,J.Hu,T.Ye,Z.Gou,S.Ma,andX.Zhang. Deepseekmath-v2: Towardsself-verifiablemathematicalreasoning,2025. URLhttps://arxiv.org/abs/25 11.22570. N.Shazeer. Fasttransformerdecoding: Onewrite-headisallyouneed. CoRR,abs/1911.02150, 2019. URLhttp://arxiv.org/abs/1911.02150. N.Shazeer. Gluvariantsimprovetransformer. arXivpreprintarXiv:2002.05202,2020. F.Shi,M.Suzgun,M.Freitag,X.Wang,S.Srivats,S.Vosoughi,H.W.Chung,Y.Tay,S.Ruder, D.Zhou,D.Das,andJ.Wei. Languagemodelsaremultilingualchain-of-thoughtreasoners. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda,May1-5,2023.OpenReview.net,2023. URLhttps://openreview.net/forum?i d=fR3wGCk-IXp. J.Su,M.Ahmed,Y.Lu,S.Pan,W.Bo,andY.Liu. Roformer: Enhancedtransformerwithrotary positionembedding. Neurocomputing,568:127063,2024. 51


M.Suzgun,N.Scales,N.Schärli,S.Gehrmann,Y.Tay,H.W.Chung,A.Chowdhery,Q.V.Le, E.H.Chi,D.Zhou,etal. Challengingbig-benchtasksandwhetherchain-of-thoughtcansolve them. arXivpreprintarXiv:2210.09261,2022. G. Tsoukalas, J. Lee, J. Jennings, J. Xin, M. Ding, M. Jennings, A. Thakur, and S. Chaudhuri. Putnambench: Evaluatingneuraltheorem-proversontheputnammathematicalcompetition, 2024. URLhttps://arxiv.org/abs/2407.11214. A.Vaswani,N.Shazeer,N.Parmar,J.Uszkoreit,L.Jones,A.N.Gomez,Ł.Kaiser,andI.Polo- sukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. L.Wang,H.Gao,C.Zhao,X.Sun,andD.Dai. Auxiliary-loss-freeloadbalancingstrategyfor mixture-of-experts. CoRR,abs/2408.15664,2024a. URLhttps://doi.org/10.48550/arX iv.2408.15664. L.Wang, Y.Cheng, Y.Shi, Z.Mo, Z.Tang, W.Xie, T.Wu, L.Ma, Y.Xia, J.Xue, etal. Tilelang: Bridge programmability and performance in modern neural kernels. In The Fourteenth InternationalConferenceonLearningRepresentations,2026. Y.Wang,X.Ma,G.Zhang,Y.Ni,A.Chandra,S.Guo,W.Ren,A.Arulraj,X.He,Z.Jiang,T.Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen. Mmlu-pro: A more robust and challengingmulti-tasklanguageunderstandingbenchmark. CoRR,abs/2406.01574,2024b. URLhttps://doi.org/10.48550/arXiv.2406.01574. J.Wei,Z.Sun,S.Papay,S.McKinney,J.Han,I.Fulford,H.W.Chung,A.T.Passos,W.Fedus, andA.Glaese. Browsecomp: Asimpleyetchallengingbenchmarkforbrowsingagents. arXiv preprintarXiv:2504.12516,2025. T.Wei,J.Luan,W.Liu,S.Dong,andB.Wang. Cmath: Canyourlanguagemodelpasschinese elementaryschoolmathtest?,2023. G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis. Efficient streaming language models with attentionsinks. InTheTwelfthInternationalConferenceonLearningRepresentations,ICLR 2024,Vienna,Austria,May7-11,2024.OpenReview.net,2024. URLhttps://openreview .net/forum?id=NG7sS51zVF. Z.Xie,Y.Wei,H.Cao,C.Zhao,C.Deng,J.Li,D.Dai,H.Gao,J.Chang,K.Yu,L.Zhao,S.Zhou, Z.Xu,Z.Zhang,W.Zeng,S.Hu,Y.Wang,J.Yuan,L.Wang,andW.Liang. mhc: Manifold- constrainedhyper-connections,2026. URLhttps://arxiv.org/abs/2512.24880. L.Xu,H.Hu,X.Zhang,L.Li,C.Cao,Y.Li,Y.Xu,K.Sun,D.Yu,C.Yu,Y.Tian,Q.Dong,W.Liu, B.Shi,Y.Cui,J.Li,J.Zeng,R.Wang,W.Xie,Y.Li,Y.Patterson,Z.Tian,Y.Zhang,H.Zhou, S.Liu,Z.Zhao,Q.Zhao,C.Yue,X.Zhang,Z.Yang,K.Richardson,andZ.Lan. CLUE:Achi- neselanguageunderstandingevaluationbenchmark. InD.Scott,N.Bel,andC.Zong,editors, Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 4762–4772. International Com- mitteeonComputationalLinguistics,2020. doi: 10.18653/V1/2020.COLING-MAIN.419. URL https://doi.org/10.18653/v1/2020.coling-main.419. J.Yang,K.Lieret,C.E.Jimenez,A.Wettig,K.Khandpur,Y.Zhang,B.Hui,O.Press,L.Schmidt, andD.Yang. Swe-smith: Scalingdataforsoftwareengineeringagents, 2025. URLhttps: //arxiv.org/abs/2504.21798. 52


R.Zellers,A.Holtzman,Y.Bisk,A.Farhadi,andY.Choi. HellaSwag: Canamachinereallyfinish yoursentence? InA.Korhonen,D.R.Traum,andL.Màrquez,editors,Proceedingsofthe57th ConferenceoftheAssociationforComputationalLinguistics,ACL2019,Florence,Italy,July 28-August2,2019,Volume1: LongPapers,pages4791–4800.AssociationforComputational Linguistics,2019. doi: 10.18653/v1/p19-1472. URLhttps://doi.org/10.18653/v1/p1 9-1472. C. Zhang, K. Du, S. Liu, W. Kwon, X. Mo, Y. Wang, X. Liu, K. You, Z. Li, M. Long, J. Zhai, J. Gonzalez, and I. Stoica. Jenga: Effective memory management for serving llm with heterogeneity. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP ’25, page 446–461, New York, NY, USA, 2025a. Association for Computing Machinery. ISBN 9798400718700. doi: 10.1145/3731569.3764823. URL https://doi.org/10.1145/3731569.3764823. S.Zhang,N.Zheng,H.Lin,Z.Jiang,W.Bao,C.Jiang,Q.Hou,W.Cui,S.Zheng,L.-W.Chang, Q. Chen, and X. Liu. Comet: Fine-grained computation-communication overlapping for mixture-of-experts. 2025b. URLhttps://arxiv.org/abs/2502.19811. C.Zhao,L.Zhao,J.Li,Z.Xu,andC.Xu. Deepgemm: cleanandefficientfp8gemmkernelswith fine-grainedscaling. https://github.com/deepseek-ai/DeepGEMM,2025. W.Zhong,R.Cui,Y.Guo,Y.Liang,S.Lu,Y.Wang,A.Saied,W.Chen,andN.Duan. AGIEval: A human-centricbenchmarkforevaluatingfoundationmodels. CoRR,abs/2304.06364,2023. doi: 10.48550/arXiv.2304.06364. URLhttps://doi.org/10.48550/arXiv.2304.06364. D. Zhu, H. Huang, Z. Huang, Y. Zeng, Y. Mao, B. Wu, Q. Min, and X. Zhou. Hyper- connections. InTheThirteenthInternationalConferenceonLearningRepresentations,ICLR 2025,Singapore,April24-28,2025.OpenReview.net,2025. URLhttps://openreview.net /forum?id=9FqARW7dwB. X.Zhu,D.Cheng,H.Li,K.Zhang,E.Hua,X.Lv,N.Ding,Z.Lin,Z.Zheng,andB.Zhou. How tosynthesizetextdatawithoutmodelcollapse? arXivpreprintarXiv:2412.14689,2024. T.Y.Zhuo,M.C.Vu,J.Chim,H.Hu,W.Yu,R.Widyasari,I.N.B.Yusuf,H.Zhan,J.He,I.Paul, S. Brunner, C. Gong, J. Hoang, A. R. Zebaze, X. Hong, W. Li, J. Kaddour, M. Xu, Z. Zhang, P. Yadav, and et al. Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions. In The Thirteenth International Conference on Learning Representations,ICLR2025,Singapore,April24-28,2025.OpenReview.net,2025. URLhttp s://openreview.net/forum?id=YrycTjllL0. 53


Appendix A. Author List and Acknowledgment A.1. AuthorList Authorsarelistedalphabeticallybytheirfirstname. Namesmarkedwithdenoteindividuals whohavedepartedfromourteam. Research & Engineering: Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, BochaoWu,BoweiZhang,ChaofanLin,ChenDong,ChengdaLu,ChenggangZhao,Chengqi Deng,ChenhaoXu,ChenzeShao,ChongRuan,ConnerSun,DamaiDai,DayaGuo,Dejian Yang, Deli Chen, Donghao Li, Erhang Li, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai,GuangboHao,GuantingChen,GuoaiCao,GuolaiMeng,GuoweiLi,HanYu,HanZhang, HanweiXu,HaoLi,HaofenLiang,HaolingZhang,HaomingLuo,HaoranWei,HaotianYuan, HaoweiZhang,HaowenLuo,HaoyuChen,HaozheJi,HonghuiDing,HongxuanTang,Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J. Yang, J.Q. Zhu, Jia Yu, Jialiang Huang, Jiasheng Ye, JiashiLi,JiaxinXu,JiewenHu,JinYan,JingchangChen,JingliZhou,JingtingXiang,Jingyang Yuan,JingyuanCheng,JinhuaZhu,JipingYu,JosephSun,JunRan,JunguangJiang,JunjieQiu, JunlongLi,JunxiaoSong,KaiDong,KaigeGao,KangGuan,KexingZhou,KezhaoHuang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Li Zhang, Liang Zhao, Lihua Guo, Lingxiao Luo, Linwang Ma, Litong Wang, Liyu Cai, Liyue Zhang, Longhao Chen, M.S. Di, M.Y Xu, MaxMei,MingchuanZhang,MinghuaZhang,MinghuiTang,MingxuZhou,PanpanHuang, PeixinCong,PeiyiWang,QianchengWang,QihaoZhu,QingyangLi,QinyuChen,QiushiDu, QiweiJiang,RuiTian,RuifanXu,RuijieLu,RuilingXu,RuiqiGe,RuisongZhang,RuizhePan, RunjiWang,RunqianChen,RunqiuYin,RunxinXu,RuomengShen,RuoyuZhang,S.H.Liu, ShanghaoLu,ShangyanZhou,ShanhuangChen,ShaofeiCai,ShaohengNie,ShaoyuanChen, ShengdingHu,ShengyuLiu,ShiqiangHu,ShirongMa,ShiyuWang,ShuipingYu,Shunfeng Zhou,ShutingPan,ShuyingYu,SongyangZhou,TaoNi,TaoYun,TianJin,TianPei,TianYe, TianleLin,TianranJi,TianyiCui,TianyuanYue,TingtingYu,TunWang,W.Zhang,Wangding Zeng,WeilinZhao,WenLiu,WenfengLiang,WenjiePang,WenjingLuo,WenjingYao,Wenjun Gao,WenkaiYang,WenlveHuang,WentaoZhang,WentingMa,XiGao,XiangHe,Xiangwen Wang,XiaoBi,XiaodongLiu,XiaohanWang,XiaokangChen,XiaokangZhang,XiaotaoNie, XinCheng,XinLiu,XinXie,XingchaoLiu,XingchenLiu,XingkaiYu,XingyouLi,XinyuYang, XuChen,XuanyuWang,XuechengSu,XuhengLin,XuweiFu,Y.C.Yan,Y.Q.Wang,Y.W.Ma, YanfengLuo,YangZhang,YanhongXu,YanruMa,YanwenHuang,YaoLi,YaoLi,YaoZhao, YaofengSun,YaohuiWang,YiQian,YiYu,YichaoZhang,YifanDing,YifanShi,YijiaWu,Yiliang Xiong,YingHe,YingZhou,YingjiaLuo,YinminZhong,YishiPiao,YisongWang,YixiangZhang, YixiaoChen,YixuanTan,YixuanWei,YiyangMa,YiyuanLiu,YonglunYang,YongqiangGuo, YongtongWu,YuWu,YuanCheng,YuanOu,YuanfanXu,YuanhaoLi,YuduanWang,Yuhan Wu,YuhaoMeng,YuhengZou,YuKunLi,YunfanXiong,YupengChen,YuqianCao,Yuqian Wang,YushunZhang,YutongLin,YuxianGu,YuxiangLuo,YuxiangYou,YuxuanLiu,Yuxuan Zhou,YuyangZhou,YuzhenHuang,Z.F.Wu,ZehaoWang,ZehuaZhao,ZehuiRen,Zhangli Sha,ZheFu,ZheanXu,ZhendaXie,ZhengyanZhang,ZhewenHao,ZhibinGou,ZhichengMa, ZhigangYan,ZhihongShao,ZhixianHuang,ZhixuanChen,ZhiyuWu,ZhizhouRen,Zhuoshu Li,ZhupingZhang,ZianXu,ZihaoWang,ZihuiGu,ZijiaZhu,ZilinLi,ZipengZhang*,Ziwei Xie,ZiyiGao,ZizhengPan,ZongqingYao. Business&Compliance: ChenchenLing,ChengyuHou,DongjieJi,FangWei,HengqingZhang, JiaLuo,JiaSong,JialuCai,JianLiang,JiangtingZhou,JieyuYang,JinChen,JingziZhou,Junmin Zheng,LeyiXia,LinyanZhu,MiaojunWang,MingmingLi,MinminHan,NingWang,Panpan 54


Wang, PengZhang, RuyiChen, ShangmianSun, ShaoqingWu, W.L.Xiao, WeiAn, Wenqing Hou, Xianzu Wang, Xiaowen Sun, Xiaoxiang Wang, Xinyu Zhang, Xueyin Chen, Yao Xu, Yi Shao,YilingMa,YingTang,YuehanYang,YuerXu,YukunZha,YupingLin,YutingYan,Zekai Zhang,ZheJu,ZherenGao,ZhongyuWu,ZihuaQu,ZiyiWan. A.2. Acknowledgment WewouldliketothankDollyDengandothertestersfortheirvaluablesuggestionsandfeedback regardingthecapabilitiesofDeepSeek-V4seriesmodels. B. Evaluation Details Table9 | AgenticSearchvs. RetrievalAugmentedSearchforDeepSeek-V4-Pro. Difficulty Category # AgentWin RAGWin Tie Agent% RAG% Tie% ObjectiveQ&A(客观问答) 196 110 43 43 56.1 21.9 21.9 Easy SubjectiveQ&A(主观问答) 321 198 56 67 61.7 17.4 20.9 ObjectiveQ&A(客观问答) 168 102 33 33 60.7 19.6 19.6 Hard SubjectiveQ&A(主观问答) 184 126 27 31 68.5 14.7 16.8 Total(总计) 869 536 159 174 61.7 18.3 20.0 Table 10 | Cost Comparison:Agentic Search vs. Retrieval Augmented Search (Mean) for DeepSeek-V4-Pro. MostofthetoolcallsareparallelforAgenticSearch. Version ToolCalls Prefill(tokens) Output(tokens) V4AgenticSearch 16.2 13649 1526 V4RetrievalAugmentedSearch — 10453 1308 Table 11 | Comparative Evaluation of DeepSeek-V4-Pro and DeepSeek-V3.2 on Search Q&A Tasks. InternalEvaluation(内部综合评估) Category Subcategory # V4win V3.2win tie V4% V3.2% tie% Single-valueSearch(单值信息查找) 95 36 10 49 37.9 10.5 51.6 Objective EntitySearch(实体信息查找) 99 24 7 68 24.2 7.1 68.7 Q&A EnumerativeSearch(枚举型信息查找) 95 19 8 68 20.0 8.4 71.6 (客观问答) Subtotal(小计) 289 79 25 185 27.3 8.7 64.0 CausalAnalysis(原因分析) 100 28 5 67 28.0 5.0 67.0 Comparison(对比) 96 28 20 48 29.2 20.8 50.0 AdviceSeeking(寻求建议) 92 23 8 61 25.0 8.7 66.3 Subjective Recommendation(推荐) 95 26 19 50 27.4 20.0 52.6 Q&A Planning&Strategy(攻略计划) 92 32 11 49 34.8 12.0 53.3 (主观问答) Opinion&Evaluation(评价看法) 96 30 8 58 31.2 8.3 60.4 TrendAnalysis(趋势分析) 96 23 3 70 24.0 3.1 72.9 Subtotal(小计) 667 190 74 403 28.5 11.1 60.4 TOTAL(总计) 956 269 99 588 28.1 10.4 61.5 55


Figure14 | Exampleoutputofataskthatrequirescomparingtworegularinvestmentstrategies fortheNASDAQ. Figure15 | Exampleoutputofataskwhichrequiresresearching2020-2025NobelSciencePrizes andgeneratingananalyticalPDFreport. 56


Table12 | ComparativeAnalysisofDeepSeek-V4-ProandGemini-3.1-ProinChineseFunctional Writing. InternalEvaluation(内部综合评估) Category Subcategory # DSwin Gemwin Tie DS% Gem% Tie% Report(报告) 527 350 162 15 66.41 30.74 2.85 Proposal(方案策划) 291 181 103 7 62.20 35.40 2.41 Education(教育培训) 159 100 56 3 62.89 35.22 1.89 Email&Letter(邮件书信) 146 107 37 2 73.29 25.34 1.37 Business Writing Notice(通知公告) 72 43 24 5 59.72 33.33 6.94 (办公文本) Professional(专业文本) 63 34 27 2 53.97 42.86 3.17 Recruitment(招聘求职) 42 27 15 0 64.29 35.71 0.00 Technical(技术文本) 29 22 7 0 75.86 24.14 0.00 Review(介绍评价) 20 15 5 0 75.00 25.00 0.00 Subtotal(小计) 1349 879 436 34 65.16 32.32 2.52 SocialMedia(社交媒体文案) 267 156 101 10 58.43 37.83 3.75 AdCopy(广告商品文案) 214 109 98 7 50.93 45.79 3.27 Long-formContent(内容平台长文) 99 71 25 3 71.72 25.25 3.03 Media NewsReport(新闻报道) 51 27 22 2 52.94 43.14 3.92 Writing Advertorial(营销软文) 17 12 4 1 70.59 23.53 5.88 (媒体文本) Headline(标题) 11 7 4 0 63.64 36.36 0.00 NarrationScript(口播文案) 4 2 1 1 50.00 25.00 25.00 Comment(评论) 3 2 1 0 66.67 33.33 0.00 Subtotal(小计) 666 386 256 24 57.96 38.44 3.60 Congratulatory(祝贺文本) 101 54 41 6 53.47 40.59 5.94 Communication(沟通回复) 100 71 26 3 71.00 26.00 3.00 Everyday Reflection(心得感想) 90 68 17 5 75.56 18.89 5.56 Writing Review(介绍评价) 55 44 9 2 80.00 16.36 3.64 (生活文本) Comment(评论) 44 34 8 2 77.27 18.18 4.55 Subtotal(小计) 390 271 101 18 69.49 25.90 4.62 Speech(发言稿) 226 135 85 6 59.73 37.61 2.65 NarrationScript(口播文案) 51 25 23 3 49.02 45.10 5.88 Oral SalesScript(话术) 31 22 6 3 70.97 19.35 9.68 Writing (口头文本) Dialogue(对话文本) 10 4 6 0 40.00 60.00 0.00 Congratulatory(祝贺文本) 1 1 0 0 100.00 0.00 0.00 Subtotal(小计) 319 187 120 12 58.62 37.62 3.76 AdministrativeDoc(事务文书) 117 60 53 4 51.28 45.30 3.42 PersonalDoc(个人文书) 73 45 27 1 61.64 36.99 1.37 Official GovernmentDoc(行政公文) 34 19 14 1 55.88 41.18 2.94 Document (公文文本) Speech(发言稿) 3 1 2 0 33.33 66.67 0.00 EssayWriting(申论写作) 3 1 1 1 33.33 33.33 33.33 Subtotal(小计) 230 126 97 7 54.78 42.17 3.04 ResearchPaper(学术论文) 104 67 32 5 64.42 30.77 4.81 Academic Coursework(课程作业) 90 53 35 2 58.89 38.89 2.22 Writing AcademicSupport(学术辅助) 15 11 3 1 73.33 20.00 6.67 (学术文本) ScienceOutreach(专业科普) 7 6 1 0 85.71 14.29 0.00 Subtotal(小计) 216 137 71 8 63.43 32.87 3.70 Total(总计) 3170 1986 1081 103 62.65 34.10 3.25 57


Table13 | ComparativeAnalysisofDeepSeek-V4-ProandGemini-3.1-ProinChineseCreative Writing. InstructionFollowing(指令遵循) WritingQuality(写作质量) Subcategory(文体) # DS Gem Tie DS% Gem% Tie% DS Gem Tie DS% Gem% Tie% Fiction(小说故事) 836 504 323 5 60.58 38.82 0.60 672 157 3 80.77 18.87 0.36 GeneralFiction(泛小说故事) 662 368 290 3 55.67 43.87 0.45 467 194 0 70.65 29.35 0.00 FanFiction(同人文) 410 253 150 3 62.32 36.95 0.74 338 67 1 83.25 16.50 0.25 GeneralFanFic.(泛同人文) 202 111 90 1 54.95 44.55 0.50 161 40 1 79.70 19.80 0.50 Narrative(记叙文) 171 115 54 2 67.25 31.58 1.17 141 30 0 82.46 17.54 0.00 GeneralProse(泛散文) 124 83 40 1 66.94 32.26 0.81 88 36 0 70.97 29.03 0.00 Prose(散文) 112 74 38 0 66.07 33.93 0.00 92 20 0 82.14 17.86 0.00 WritingStyle(文笔) 112 81 31 0 72.32 27.68 0.00 86 26 0 76.79 23.21 0.00 ClassicalPoetry(古诗文) 48 24 24 0 50.00 50.00 0.00 39 9 0 81.25 18.75 0.00 ModernPoetry(现代诗) 43 23 20 0 53.49 46.51 0.00 32 11 0 74.42 25.58 0.00 Lyrics(歌词) 30 8 22 0 26.67 73.33 0.00 16 14 0 53.33 46.67 0.00 LiteraryAppreciation(赏析) 27 20 7 0 74.07 25.93 0.00 18 9 0 66.67 33.33 0.00 GeneralArgument.(泛议论文) 24 15 9 0 62.50 37.50 0.00 17 7 0 70.83 29.17 0.00 GeneralNarrative(泛记叙文) 23 11 12 0 47.83 52.17 0.00 15 8 0 65.22 34.78 0.00 GeneralClassical(泛古文诗歌) 9 5 4 0 55.56 44.44 0.00 5 4 0 55.56 44.44 0.00 CreativeWriting(创意写作) 6 2 4 0 33.33 66.67 0.00 4 2 0 66.67 33.33 0.00 Argumentative(议论文) 5 5 0 0 100.00 0.00 0.00 5 0 0 100.00 0.00 0.00 GeneralMod.Poetry(泛现代诗) 2 1 1 0 50.00 50.00 0.00 2 0 0 100.00 0.00 0.00 Total(总计) 2837 1703 1119 15 60.03 39.44 0.53 2198 634 5 77.48 22.35 0.18 Table14 | DeepSeek-V4-Provs. Claude-Opus-4.5onComplexInstructionFollowingandMulti- TurnWriting. InternalEvaluation(内部综合评估) Category # DS Opus Tie DS% Opus% Tie% ComplexInst. Following(复杂指令跟随) 49 23 26 0 46.9% 53.1% 0.0% Multi-TurnWriting(多轮写作) 147 67 76 4 45.6% 51.7% 2.7% Total(总计) 196 90 102 4 45.9% 52.0% 2.0% 58


Vissza a tetejére