DeepSeek-V4:
Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-AI
research@deepseek.com
Abstract
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-
Experts(MoE)languagemodels—DeepSeek-V4-Prowith1.6Tparameters(49Bactivated)and
DeepSeek-V4-Flashwith284Bparameters(13Bactivated)—bothsupportingacontextlengthof
onemilliontokens. DeepSeek-V4seriesincorporateseveralkeyupgradesinarchitectureandop-
timization: (1)ahybridattentionarchitecturethatcombinesCompressedSparseAttention(CSA)
andHeavilyCompressedAttention(HCA)toimprovelong-contextefficiency;(2)Manifold-
Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3)
and the Muon optimizer for faster convergence and greater training stability. We pre-train
bothmodelsonmorethan32Tdiverseandhigh-qualitytokens,followedbyacomprehensive
post-trainingpipelinethatunlocksandfurtherenhancestheircapabilities. DeepSeek-V4-Pro-
Max,themaximumreasoningeffortmodeofDeepSeek-V4-Pro,redefinesthestate-of-the-artfor
openmodels,outperformingitspredecessorsincoretasks. Meanwhile,DeepSeek-V4seriesare
highlyefficientinlong-contextscenarios. Intheone-million-tokencontextsetting,DeepSeek-
V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared
withDeepSeek-V3.2. Thisenablesustoroutinelysupportone-million-tokencontexts,thereby
makinglong-horizontasksandfurthertest-timescalingmorefeasible. Themodelcheckpoints
areavailableathttps://huggingface.co/collections/deepseek-ai/deepseek-v4.
100
80
60
40
20
0
SimpleQA HLE Apex Codeforces SWE Terminal Toolathlon
Verified (Pass@1) Shortlist (Rating) Verified Bench 2.0 (Pass@1)
(Pass@1) (Pass@1) (Resolved) (Acc)
)%(
1@ssaP
/ ycaruccA
1.2
1.0
DeepSeek-V4-Pro-Max Claude-Opus-4.6-Max GPT-5.4-xHigh Gemini-3.1-Pro-High 0.8
90.2 85.9 89.1 32063168 3052 0.6
80.680.880.6 75.6 78.1 75.1 0.4
67.9 68.5
65.4 0.2
57.9
51.8 54.6 0.0
46.245.3 44.4 47.2 48.8 0 25 T 6 oken Po 51 s 2 ition (K) 768 1024
40.039.8 37.7
Knowledge & Reasoning Agentic Capabilities
)T(
sPOLF
nekoT-elgniS
DeepSeek-V3.2
DeepSeek-V4-Pro
DeepSeek-V4-Flash
3.7× lower
9.8× lower
50
40
30
20
10
0
0 256 512 768 1024
Sequence Length (K)
)BG(
ehcaC
VK
detalumuccA
DeepSeek-V3.2
DeepSeek-V4-Pro
DeepSeek-V4-Flash
9.5× smaller
13.7× smaller
Figure1 | Left: benchmarkperformanceofDeepSeek-V4-Pro-Maxanditscounterparts. Right:
inferenceFLOPsandKVcachesizeofDeepSeek-V4seriesandDeepSeek-V3.2.

---

Contents
1 Introduction 4
2 Architecture 6
2.1 DesignsInheritedfromDeepSeek-V3. . . . . . . . . . . . . . . . . . . . . . . . . . 7
2.2 Manifold-ConstrainedHyper-Connections . . . . . . . . . . . . . . . . . . . . . . 7
2.3 HybridAttentionwithCSAandHCA . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.3.1 CompressedSparseAttention . . . . . . . . . . . . . . . . . . . . . . . . . . 9
2.3.2 HeavilyCompressedAttention . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.3.3 OtherDetails . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
2.3.4 EfficiencyDiscussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.4 MuonOptimizer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
3 GeneralInfrastructures 15
3.1 Fine-GrainedCommunication-ComputationOverlapinExpertParallelism . . . . 15
3.2 FlexibleandEfficientKernelDevelopmentwithTileLang . . . . . . . . . . . . . . 16
3.3 High-PerformanceBatch-InvariantandDeterministicKernelLibraries . . . . . . 18
3.4 FP4Quantization-AwareTraining. . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
3.5 TrainingFramework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
3.5.1 EfficientImplementationofMuon . . . . . . . . . . . . . . . . . . . . . . . 20
3.5.2 Cost-EffectiveandMemory-EfficientImplementationofmHC . . . . . . . 21
3.5.3 ContextualParallelismforLong-ContextAttention . . . . . . . . . . . . . 21
3.5.4 ExtendedAutomaticDifferentiationforFlexibleActivationCheckpointing 21
3.6 InferenceFramework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
3.6.1 KVCacheStructureandManagement . . . . . . . . . . . . . . . . . . . . . 22
3.6.2 On-DiskKVCacheStorage . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
4 Pre-Training 24
4.1 DataConstruction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24
4.2 Pre-TrainingSetups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
4.2.1 ModelSetups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
4.2.2 TrainingSetups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
4.2.3 MitigatingTrainingInstability . . . . . . . . . . . . . . . . . . . . . . . . . 26
4.3 Evaluations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
4.3.1 EvaluationBenchmarks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
4.3.2 EvaluationResults . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
2

---

5 Post-Training 29
5.1 Post-TrainingPipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
5.1.1 SpecialistTraining . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
5.1.2 On-PolicyDistillation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
5.2 RLandOPDInfrastructures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
5.2.1 FP4QuantizationIntegration . . . . . . . . . . . . . . . . . . . . . . . . . . 34
5.2.2 EfficientTeacherSchedulingforFull-VocabularyOPD . . . . . . . . . . . 34
5.2.3 PreemptibleandFault-TolerantRolloutService . . . . . . . . . . . . . . . 34
5.2.4 ScalingRLFrameworkforMillion-TokenContext . . . . . . . . . . . . . . 35
5.2.5 SandboxInfrastructureforAgenticAI . . . . . . . . . . . . . . . . . . . . . 35
5.3 StandardBenchmarkEvaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
5.3.1 EvaluationSetup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36
5.3.2 EvaluationResults . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
5.4 PerformanceonReal-WorldTasks . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
5.4.1 ChineseWriting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
5.4.2 Search . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
5.4.3 White-CollarTask . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
5.4.4 CodeAgent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
6 Conclusion,Limitations,andFutureDirections 44
A AuthorListandAcknowledgment 54
A.1 AuthorList . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54
A.2 Acknowledgment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
B EvaluationDetails 55
3

---

1. Introduction
The emergence of reasoning models (DeepSeek-AI, 2025; OpenAI, 2024c) has established a
newparadigmoftest-timescaling,drivingsubstantialperformancegainsforLargeLanguage
Models(LLMs). However,thisscalingparadigmisfundamentallyconstrainedbythequadratic
computational complexity of the vanilla attention mechanism (Vaswani et al., 2017), which
createsaprohibitivebottleneckforultra-longcontextsandreasoningprocesses. Concurrently,
the emergence of long-horizon scenarios and tasks — from complex agentic workflows to
massive cross-document analysis — has also made efficient support for ultra-long contexts
critical for future progress. While recent open-source efforts (Bai et al., 2025a; DeepSeek-AI,
2024;MiniMax,2025;Qwen,2025)haveadvancedgeneralcapabilities,thiscorearchitectural
inefficiencyinhandlingultra-longsequencesremainsakeyimpediment,limitingfurthergains
fromtest-timescalingandhinderingfurtherexplorationintolong-horizonscenariosandtasks.
Inordertobreaktheefficiencybarrierinultra-longcontexts,wedeveloptheDeepSeek-V4
series,includingthepreviewversionsofDeepSeek-V4-Prowith1.6Tparameters(49Bactivated)
andDeepSeek-V4-Flashwith284Bparameters(13Bactivated). Througharchitecturalinnova-
tions,DeepSeek-V4seriesachieveadramaticleapincomputationalefficiencyforprocessing
ultra-longsequences. Thisbreakthroughenablesefficientsupportforacontextlengthofone
milliontokens,usheringinaneweraofmillion-lengthcontextsfornext-generationLLMs. We
believe our capability to efficiently handle ultra-long sequences unlocks the next frontier of
test-timescaling,pavesthewayfordeeperresearchintolong-horizontasks,andestablishesa
necessaryfoundationforexploringfutureparadigmslikeonlinelearning.
Compared with the DeepSeek-V3 architecture (DeepSeek-AI, 2024), DeepSeek-V4 series
retaintheDeepSeekMoEframework(Daietal.,2024)andMulti-TokenPrediction(MTP)strategy,
whileintroducingseveralkeyinnovationsinarchitectureandoptimization. Toenhancelong-
context efficiency, we design a hybrid attention mechanism combining Compressed Sparse
Attention(CSA)andHeavilyCompressedAttention(HCA).CSAcompressestheKVcaches
alongthesequencedimensionandthenperformsDeepSeekSparseAttention(DSA)(DeepSeek-
AI, 2025), whereas HCA applies more aggressive compression to the KV caches but keeps
dense attention. To strengthen modeling capability, we incorporate Manifold-Constrained
Hyper-Connections (mHC) (Xie et al., 2026) that upgrade conventional residual connections.
Additionally, we introduce the Muon (Jordan et al., 2024; Liu et al., 2025) optimizer to the
trainingofDeepSeek-V4series,leadingtofasterconvergenceandimprovedtrainingstability.
ToenableefficienttrainingandinferenceforDeepSeek-V4seriesaswellasproductivede-
velopment,weintroduceseveralinfrastructureoptimizations. First,wedesignandimplement
asinglefusedkernelforMoEmodulesthatfullyoverlapscomputation,communication,and
memoryaccess. Second,weemployTileLang(Wangetal.,2026),aDomain-SpecificLanguage
(DSL)tobalancedevelopmentproductivityandruntimeefficiency. Third,weprovideefficient
batch-invariantanddeterministickernellibrariestoensurebitwisereproducibilityacrosstrain-
ing and inference. Fourth, we incorporate FP4 quantization-aware training for MoE expert
weightsandtheindexerQKpathtoreducememoryandcomputation. Fifth,forthetraining
framework,weextendtheautogradframeworkwithtensor-levelcheckpointingforfine-grained
recomputationcontrol;andweenhancetrainingefficiencywithahybridZeROstrategyforthe
Muonoptimizer,cost-effectivemHCimplementationsviarecomputationandfusedkernels,and
two-stage contextual parallelism to manage compressed attention. Finally, for the inference
framework,wedesignaheterogeneousKVcachestructurewithon-diskstoragestrategiesto
enableefficientshared-prefixreuse.
4

---

ByemployinghybridCSAandHCA,alongwithprecisionoptimizationsoncomputation
andstorage,DeepSeek-V4seriesachievesignificantlylowerinferenceFLOPsandasubstantially
reducedKVcachesizecomparedwithDeepSeek-V3.2,especiallyinlong-contextsettings. The
rightpartofFigure1demonstratestheestimatedsingle-tokeninferenceFLOPsandaccumulated
KVcachesizeofDeepSeek-V3.2andDeepSeek-V4series. Inthescenarioof1M-tokencontext,
evenDeepSeek-V4-Pro,whichhasalargernumberofactivatedparameters,attainsonly27%
of the single-token FLOPs (measured in equivalent FP8 FLOPs) and 10% of the KV cache
sizerelativetoDeepSeek-V3.2. Furthermore,DeepSeek-V4-Flash,withitssmallernumberof
activatedparameters,pushesefficiencyevenfurther: inthe1M-tokencontextsetting,itachieves
only10%ofthesingle-tokenFLOPsand7%oftheKVcachesizecomparedwithDeepSeek-V3.2.
Additionally,forDeepSeek-V4series,theroutedexpertparametersutilizeFP4precision. While
the peak FLOPs for FP4 × FP8 operations are currently the same as FP8 × FP8 on existing
hardware,theycantheoreticallybeimplementedtobe1/3moreefficientonfuturehardware,
whichwillfurtherenhancetheefficiencyofDeepSeek-V4series.
Duringpre-training,wetrainDeepSeek-V4-Flashon32TtokensandDeepSeek-V4-Proon33T
tokens,respectively. Afterpre-training,thesetwomodelscannativelyandefficientlysupport
1M-length contexts. In our internal evaluations, DeepSeek-V4-Flash-Base already surpasses
DeepSeek-V3.2-Baseacrossamajorityofbenchmarkswithitsmoreparameter-efficientdesign.
DeepSeek-V4-Pro-Basefurtherextendsthisadvantagetosetanewperformancestandardamong
DeepSeekfoundationmodels,achievingcomprehensivesuperiorityacrossreasoning,coding,
long-context,andworldknowledgetasks.
Thepost-trainingpipelineofDeepSeek-V4seriesfeaturesatwo-stageparadigm: theinde-
pendentcultivationofdomain-specificexperts,followedbyunifiedmodelconsolidationvia
on-policydistillation(LuandLab,2025). Initially,foreachtargetdomain—suchasmathematics,
coding,agent,andinstructionfollowing—aseparateexpertmodelistrainedindependently.
ThebasemodelfirstundergoesSupervisedFine-Tuning(SFT)onhigh-quality,domain-specific
datatoestablishfoundationalcapabilities. Subsequently,ReinforcementLearning(RL)isap-
pliedusingGroupRelativePolicyOptimization(GRPO)(DeepSeek-AI,2025),whichfurther
optimizesthemodelfordomain-alignedbehaviorsguidedbyrewardmodelstailoredtospecific
success criteria. This phase yields a diverse set of specialized experts, each excelling in its
respectivefield. Finally,tointegratethesedistinctproficiencies,asingleunifiedmodelistrained
throughon-policydistillation,whereintheunifiedmodelactsasthestudentlearningtooptimize
thereverseKLlosswithteachermodels.
SummaryofCoreEvaluationResults
• Knowledge: Inassessmentsofbroadworldknowledge,DeepSeek-V4-Pro-Max,themaxi-
mumreasoningeffortmodeofDeepSeek-V4-Pro,significantlyoutperformsleadingopen-
sourcemodelsontheSimpleQA(OpenAI,2024d)andChinese-SimpleQA(Heetal.,2024)
benchmarks. Regardingeducationalknowledge—evaluatedviaMMLU-Pro(Wangetal.,
2024b),HLE(Phanetal.,2025),andGPQA(Reinetal.,2023)—DeepSeek-V4-Pro-Max
shows a marginal lead over its open-source counterparts. DeepSeek-V4-Pro-Max has
significantlyclosedthegapwiththeleadingproprietarymodel,Gemini-3.1-Pro,despite
stilltrailingitintheseknowledge-basedevaluations.
• Reasoning: Throughtheexpansionofreasoningtokens,DeepSeek-V4-Pro-Maxdemon-
stratessuperiorperformancerelativetoGPT-5.2andGemini-3.0-Proonstandardreasoning
benchmarks. Nevertheless,itsperformancefallsmarginallyshortofGPT-5.4andGemini-
3.1-Pro,suggestingadevelopmentaltrajectorythattrailsstate-of-the-artfrontiermodelsby
approximately3to6months. Furthermore,DeepSeek-V4-Flash-Maxachievescomparable
5

---

MTP Modules MTP Loss
Prediction Head LM Loss
Transformer Block ×
𝐿𝐿
Post-Block Mixing
DeepSeekMoE
Residual Mixing
Pre-Block Mixing
Post-Block Mixing
CSA / HCA
Residual Mixing
Pre-Block Mixing
Embedding
Input Tokens
Figure2 | OverallarchitectureofDeepSeek-V4series. WeusehybridCSA(CompressedSparse
Attention)andHCA(HeavilyCompressedAttention)forattentionlayers,DeepSeekMoEfor
feed-forwardlayers,andstrengthenconventionalresidualconnectionswithmHC.
performancetoGPT-5.2andGemini-3.0-Pro,establishingitselfasahighlycost-effective
architectureforcomplexreasoningtasks.
• Agent: Onpublicbenchmarks,DeepSeek-V4-Pro-Maxisonparwithleadingopen-source
models,suchasKimi-K2.6andGLM-5.1,butslightlyworsethanfrontierclosedmodels.
In our internal evaluation, DeepSeek-V4-Pro-Max outperforms Claude Sonnet 4.5 and
approachesthelevelofOpus4.5.
• Long-Context: DeepSeek-V4-Pro-Maxdeliversstrongresultsonsyntheticandrealuse
caseswitha1-million-tokencontextwindow,surpassingevenGemini-3.1-Proonacademic
benchmarks.
• DeepSeek-V4-Prov.s. DeepSeek-V4-Flash: DeepSeek-V4-Flash-Maxexhibitslowerper-
formance in knowledge evaluations due to its smaller parameter scale. However, it
achieves comparable results on reasoning tasks when allocated a larger thinking bud-
get. In agent evaluations, while DeepSeek-V4-Flash-Max matches the performance of
DeepSeek-V4-Pro-Maxonseveralbenchmarks,itstilltrailsitslargercounterpartonmore
complex,high-difficultytasks.
2. Architecture
Overall,DeepSeek-V4seriesretaintheTransformer(Vaswanietal.,2017)architectureandMulti-
TokenPrediction(MTP)modules(DeepSeek-AI,2024;Gloeckleetal.,2024),whileintroducing
several key upgrades over DeepSeek-V3: (1) firstly, we introduce the Manifold-Constrained
Hyper-Connections(mHC)(Xieetal.,2026)tostrengthenconventionalresidualconnections;
6

---

(2)secondly,wedesignahybridattentionarchitecture,whichgreatlyimproveslong-context
efficiencythroughCompressedSparseAttentionandHeavilyCompressedAttention. (3)thirdly,
we employ Muon (Jordan et al., 2024; Liu et al., 2025) as the optimizer. For the Mixture-of-
Experts(MoE)components,westilladopttheDeepSeekMoE(Daietal.,2024)architecture,with
onlyminoradjustmentsfromDeepSeek-V3. TheMulti-TokenPrediction(MTP)(DeepSeek-AI,
2024; Gloeckle et al., 2024; Li et al., 2024; Qi et al., 2020) configuration remains identical to
thatofDeepSeek-V3. AllotherunspecifieddetailsfollowthesettingsestablishedinDeepSeek-
V3(DeepSeek-AI,2024). Figure2illustratestheoverallarchitectureofDeepSeek-V4,andthe
detailsaredescribedbelow.
2.1. DesignsInheritedfromDeepSeek-V3
Mixture-of-Experts. AspreviousDeepSeek-seriesmodels(DeepSeek-AI,2024;DeepSeek-AI,
2024),DeepSeek-V4seriesalsoadopttheDeepSeekMoEparadigm(Daietal.,2024)forFeed-
ForwardNetworks(FFNs),whichsetsfine-grainedroutedexpertsandsharedexperts. Different
fromDeepSeek-V3,wechangetheactivationfunctionthatcomputestheaffinityscoresfrom
Sigmoid(·) into Sqrt(Softplus(·)). For load balancing, we also employ the auxiliary-loss-free
strategy(DeepSeek-AI,2024;Wangetal.,2024a),augmentedbyaslightsequence-wisebalance
lossthatpreventsextremeimbalancewithinindividualsequences. ForDeepSeek-V4,weremove
the constraint on the number of routing target nodes, and carefully redesign the parallelism
strategytomaintaintrainingefficiency. Furthermore,comparedwithDeepSeek-V3,wereplace
the dense FFN layers in the initial several Transformer blocks with MoE layers that employ
Hashrouting(Rolleretal.,2021). TheHashroutingstrategydeterminesthetargetexpertsof
eachtokenaccordingtoapredefinedhashfunctionwithregardtotheinputtokenID.
Multi-Token Prediction. As DeepSeek-V3, DeepSeek-V4 series also set MTP modules and
objectives. GiventhattheMTPstrategyhasbeenvalidatedinDeepSeek-V3,weadoptthesame
strategyforDeepSeek-V4serieswithoutmodification.
2.2. Manifold-ConstrainedHyper-Connections
AsshowninFigure2,DeepSeek-V4seriesincorporateManifold-ConstrainedHyper-Connections
(mHC)(Xieetal.,2026)tostrengthentheconventionalresidualconnectionsbetweenadjacent
Transformerblocks. ComparedwithnaiveHyper-Connections(HC)(Zhuetal.,2025),thecore
ideaofmHCistoconstraintheresidualmappingontoaspecificmanifold,andthusenhancethe
stabilityofsignalpropagationacrosslayerswhilepreservingmodelexpressivity. Thissubsection
brieflyintroducesthestandardHCanddescribeshowwedesignmHCforstabletraining.
Standard Hyper-Connections. The standard HC expands the width of the residual stream
byafactorof𝑛 hc . Specifically,theshapeoftheresidualstreamisexpandedfrom R𝑑 to R𝑛 hc ×𝑑 ,
where 𝑑 is the hidden size of the actual layer input. Let 𝑋 𝑙 = [x𝑙,1 ;...;x𝑙,𝑛
hc
]𝑇 ∈ R𝑛 hc ×𝑑 be the
residual state before the 𝑙-th layer. HC introduces three linear mappings: an input mapping
𝐴 𝑙 ∈ R1×𝑛 hc, a residual transformation 𝐵 𝑙 ∈ R𝑛 hc ×𝑛 hc, and an output mapping𝐶 𝑙 ∈ R𝑛 hc ×1. The
updateoftheresidualstateisthenformulatedas:
𝑋 𝑙+1 = 𝐵 𝑙 𝑋 𝑙 +𝐶 𝑙 F 𝑙 (𝐴 𝑙 𝑋 𝑙 ), (1)
where F 𝑙 denotesthe 𝑙-thlayer(e.g.,anMoElayer),whoseinputandoutputshapesareboth
R𝑑 . Notethattheactuallayerinput 𝐴 𝑙 𝑋 𝑙 ∈ R𝑑 isalso𝑑-dimensional,sotheexpandedresidual
7

---

widthdoesnotinfluencethedesignoftheinnerlayers. HCdecouplestheresidualwidthfrom
the actual hidden size, offering a complementary scaling axis with minimal computational
overhead,as𝑛 istypicallymuchsmallerthanthehiddensize𝑑. However,eventhoughHC
hc
has demonstrated potential in improving model performance, we find that the training will
frequentlyexhibitnumericalinstabilitywhenstackingmultiplelayers,whichhindersthescaling
ofHC.
Manifold-ConstrainedResidualMapping. ThecoreinnovationofmHCistoconstrainthe
residualmappingmatrix 𝐵 𝑙 tothemanifoldofdoublystochasticmatrices(theBirkhoffpolytope)
M,andthusenhancethestabilityofsignalpropagationacrosslayers:
𝐵 𝑙 ∈ M ≔ {𝑀 ∈ R𝑛×𝑛 | 𝑀1𝑛 =1𝑛, 1 𝑇 𝑛 𝑀 =1 𝑇 𝑛 , 𝑀 ⩾ 0}. (2)
Thisconstraintensuresthatthespectralnormofthemappingmatrix ∥𝐵 𝑙 ∥ 2 isboundedby1,so
theresidualtransformationisnon-expansive,whichincreasesthenumericalstabilityduringboth
theforwardpassandbackpropagation. Besides,thesetM isclosedundermultiplication,which
guaranteesstabilityinthescenariosofdeepstacksofmHC.Inaddition,theinputtransformation
𝐴 𝑙 and output transformation 𝐶 𝑙 are also constrained to be non-negative and bounded via a
Sigmoidfunctiontoavoidtheriskofsignalcancellation.
DynamicParameterization. Theparametersofthreelinearmappingsaredynamicallygen-
erated, which are decomposed into a dynamic (input-dependent) component and a static
(input-independent)component. Giventheinput 𝑋 𝑙 ∈ R𝑛 hc ×𝑑 , itisfirstflattenedandnormal-
ized: 𝑋ˆ 𝑙 =RMSNorm(vec(𝑋 𝑙 )) ∈ R1×𝑛 hc 𝑑 . Then,wefollowtheconventionalHCtogeneratethe
unconstrainedrawparameters 𝐴˜ 𝑙 ∈ R1×𝑛 hc, 𝐵˜ 𝑙 ∈ R𝑛 hc ×𝑛 hc,and𝐶˜ 𝑙 ∈ R𝑛 hc ×1:
𝐴˜ 𝑙 =𝛼p 𝑙 re·(𝑋ˆ 𝑙 𝑊 𝑙 pre)+𝑆 𝑙 pre , (3)
𝐵˜ 𝑙 =𝛼r 𝑙 es·Mat(𝑋ˆ 𝑙 𝑊 𝑙 res)+𝑆 𝑙 res, (4)
𝐶˜ 𝑙 =𝛼p 𝑙 ost·(𝑋ˆ 𝑙 𝑊 𝑙 post)𝑇 +𝑆 𝑙 post , (5)
where𝑊 𝑙 pre ,𝑊 𝑙 post ∈ R𝑛 hc 𝑑×𝑛 hc and𝑊 𝑙 res ∈ R𝑛 hc 𝑑×𝑛2 hc arelearnableparametersforgeneratingthe
dynamic components; Mat(·) reshapes a vector of size 1×𝑛2 into a matrix of size 𝑛 ×𝑛 ;
hc hc hc
𝑆pre ∈ R1×𝑛 hc,𝑆post ∈ R𝑛 hc ×1,and𝑆res ∈ R𝑛 hc ×𝑛 hc arelearnablestaticbiases;and𝛼pre ,𝛼res,𝛼post ∈ R
𝑙 𝑙 𝑙 𝑙 𝑙 𝑙
arelearnablegatingfactorsinitializedtosmallvalues.
ApplyingParameterConstraints. Afterobtainingtheunconstrainedrawparameters 𝐴˜ 𝑙,𝐵˜ 𝑙,𝐶˜ 𝑙,
wethenapplyconstraintsdescribedearliertothemtoenhancethenumericalstability. Tobe
specific,fortheinputandoutputmappings,weemployaSigmoidfunction𝜎(·) toensuretheir
non-negativityandboundedness:
𝐴 𝑙 =𝜎(𝐴˜ 𝑙 ), (6)
𝐶 𝑙 =2𝜎(𝐶˜ 𝑙 ). (7)
Asfortheresidualmapping 𝐵˜ 𝑙,weprojectitontothemanifoldofdoublystochasticmatricesM.
ThisisachievedbytheSinkhorn-Knoppalgorithm,whichfirstappliesanexponentialfunction
to 𝐵˜ 𝑙 toensurepositivity,getting 𝑀(0) =exp(𝐵˜ 𝑙 ),andtheniterativelyperformscolumnandrow
normalization:
𝑀(𝑡) =T 𝑟 (T 𝑐 (𝑀(𝑡−1))), (8)
whereT 𝑟 andT 𝑐 denoterowandcolumnnormalization,respectively. Thisiterationconvergesto
aconstraineddoublystochasticmatrix 𝐵 𝑙 =𝑀(𝑡 max ). Wechoose𝑡 max =20asapracticalvalue.
8

---

Shared Key-Value Multi-Query Attention
Concatenation
Lightning Indexer
Sliding Window Selected
KV Entries Compressed …
KV Entries Index Scores
Top-k Multi-Query
…
Selector Attention
Compressed Compressed
KV Entries … Indexer Keys … Indexer Queries Queries
Token-Level Token-Level
Compressor Compressor
Hidden States of KV Tokens … Hidden State of Query Token
Figure3 | CorearchitecturesofCSA.ItcompressesthenumberofKVentriesto 1 times,and
𝑚
thenapplies DeepSeekSparseAttention forfurtheracceleration. Additionally, asmallset of
slidingwindowKVentriesiscombinedwiththeselectedcompressedKVentriestoenhance
localfine-graineddependencies.
2.3. HybridAttentionwithCSAandHCA
Asthecontextlengthreachesextremescales,theattentionmechanismemergesasthedominant
computational bottleneck in a model. For DeepSeek-V4, we design two efficient attention
architectures—CompressedSparseAttention(CSA)andHeavilyCompressedAttention(HCA)
—andemploytheirinterleavedhybridconfiguration,whichsubstantiallyreducesthecompu-
tationalcostofattentioninlong-textscenarios. CSAintegratesbothcompressionandsparse
attention strategies: it first compresses the Key-Value (KV) cache of every 𝑚 tokens into one
entry,andthenappliesDeepSeekSparseAttention(DSA)(DeepSeek-AI,2025)whereeachquery
tokenattendstoonly𝑘compressedKVentries. HCAaimsforextremecompressionbyconsol-
idatingtheKVcacheofevery𝑚′ (≫ 𝑚)tokensintoasingleentry. Thehybridarchitectureof
CSAandHCAremarkablyimprovesthelong-contextefficiencyofDeepSeek-V4series,making
one-million-token context feasible in practice. This subsection describes the core techniques
ofourhybridattentionarchitecture,andwealsoprovideanopen-sourceimplementation1 to
specifymoredetailsunambiguously.
2.3.1. CompressedSparseAttention
ThecorearchitectureofCSAisillustratedinFigure3,whichfirstcompressestheKVcacheofeach
𝑚tokensintooneentry,andthenappliesDeepSeekSparseAttentionforfurtheracceleration.
CompressedKey-ValueEntries. Let 𝐻 ∈ R𝑛×𝑑 beasequenceofinputhiddenstates,where
𝑛isthesequencelengthand𝑑 isthehiddensize. CSAfirstcomputestwoseriesofKVentries
𝐶𝑎 ,𝐶𝑏 ∈ R𝑛×𝑐 andtheircorrespondingcompressionweights 𝑍𝑎 ,𝑍𝑏 ∈ R𝑛×𝑐 ,where 𝑐 isthehead
1https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/tree/main/inference
9

---

dimension:
𝐶𝑎 = 𝐻·𝑊𝑎𝐾𝑉 , 𝐶𝑏 = 𝐻·𝑊𝑏𝐾𝑉 , (9)
𝑍𝑎 = 𝐻·𝑊𝑎𝑍 , 𝑍𝑏 = 𝐻·𝑊𝑏𝑍 , (10)
where𝑊𝑎𝐾𝑉 ,𝑊𝑏𝐾𝑉 ,𝑊𝑎𝑍 ,𝑊𝑏𝑍 ∈ R𝑑×𝑐 aretrainableparameters. Next,each𝑚KVentriesin𝐶𝑎 and
𝐶𝑏
will be compressed into one entry according to their compression weights and learnable
positionalbiases 𝐵𝑎 ,𝐵𝑏 ∈ R𝑚×𝑐 ,producing𝐶Comp ∈ R 𝑚 𝑛×𝑐 . Eachcompressedentry𝐶Comp ∈ R𝑐 is
𝑖
computedby
[𝑆𝑎 ;𝑆𝑏 ] =Softmax ([𝑍𝑎 +𝐵𝑎 ;𝑍𝑏 +𝐵𝑏]), (11)
𝑚𝑖:𝑚(𝑖+1)−1 𝑚(𝑖−1):𝑚𝑖−1 row 𝑚𝑖:𝑚(𝑖+1)−1 𝑚(𝑖−1):𝑚𝑖−1
𝑚(𝑖+1)−1 𝑚𝑖−1
∑︁ ∑︁
𝐶Comp = 𝑆𝑎⊙𝐶𝑎+ 𝑆𝑏⊙𝐶𝑏 , (12)
𝑖 𝑗 𝑗 𝑗 𝑗
𝑗=𝑚𝑖 𝑗=𝑚(𝑖−1)
where ⊙ denotestheHadamardproduct; Softmax (·) denotesthesoftmaxoperationalong
row
therowdimension,whichperformsnormalizationacrossthetotalof2𝑚elementsfromboth
𝑍𝑎 and 𝑍𝑏 . When 𝑖 =0, 𝑍𝑏 ispaddedwithnegativeinfinityand𝐶𝑏 ispadded
𝑚(𝑖−1):𝑚𝑖−1 𝑚(𝑖−1):𝑚𝑖−1
withzeros.
Notethateach𝐶Comp isderivedfrom2𝑚KVentries,buttheindexesof𝐶𝑏
usedfor
𝑖
𝐶Comp andtheindexesof𝐶𝑎 usedfor𝐶Comp
areoverlapped. Therefore,CSAinfactcompresses
𝑖 𝑖−1
thesequencelengthto 1 times.
𝑚
LightningIndexerforSparseSelection. AfterobtainingthecompressedKVentries𝐶Comp,
CSAappliestheDSAstrategytoselecttop-kcompressedKVentriesforcoreattention. First,
CSAperformsthesamecompressionoperationusedfor𝐶Comp togetcompressedindexerkeys
𝐾IComp ∈ R 𝑚 𝑛×𝑐𝐼 ,where𝑐𝐼 istheindexerheaddimension. Then,foraquerytoken𝑡,weproduce
theindexerqueries {q 𝐼 ;q 𝐼 ;...;q 𝐼 } inalow-rankmanner:
𝑡,1 𝑡,2 𝑡,𝑛𝐼
ℎ
c
𝑡
𝑄 =h𝑡 ·𝑊𝐷𝑄 , (13)
[q 𝐼 ;q 𝐼 ;...;q 𝐼 ] =q 𝐼 =c 𝑄 ·𝑊𝐼𝑈𝑄 , (14)
𝑡,1 𝑡,2 𝑡,𝑛𝐼 𝑡 𝑡
ℎ
where h𝑡 ∈ R𝑑 is the input hidden state of the query token 𝑡; c
𝑡
𝑄 ∈ R𝑑𝑐 is the compressed
latentvectorforqueries;𝑑 𝑐 denotesthequerycompressiondimension;𝑛 ℎ 𝐼 denotesthenumber
of indexer query heads; 𝑊𝐷𝑄 ∈ R𝑑×𝑑𝑐 and 𝑊𝐼𝑈𝑄 ∈ R𝑑𝑐×𝑐𝐼𝑛 ℎ 𝐼 are the down-projection and up-
projectionmatricesforindexerqueries,respectively. Next,theindexscore 𝐼 𝑡,𝑠 ∈ R betweenthe
querytoken𝑡 andaprecedingcompressedblock𝑠(𝑠<Floor( 𝑡 ))iscomputedby
𝑚
[𝑤
𝑡
𝐼
,1
;𝑤
𝑡
𝐼
,2
;...;𝑤
𝑡
𝐼
,𝑛𝐼
] =w
𝑡
𝐼 =h𝑡 ·𝑊𝑤 , (15)
ℎ
𝑛𝐼
∑︁ℎ (cid:16) (cid:17)
𝐼 𝑡,𝑠 = 𝑤 𝑡 𝐼 ,ℎ ·ReLU q 𝑡 𝐼 ,ℎ ·𝐾 𝑠 IComp , (16)
ℎ=1
where𝑊𝑤 ∈ R𝑑×𝑛 ℎ 𝐼 isalearnablematrix;𝑤𝐼 ∈ R istheweightoftheℎ-thindexerhead. Fora
𝑡,ℎ
querytoken𝑡,givenitsindexscores 𝐼 𝑡,: ,weemployatop-kselectortoselectivelyretainasubset
ofcompressedKVentries
CSprsComp
forsubsequentcoreattention:
𝑡
(cid:110) (cid:12) (cid:111)
C 𝑡 SprsComp = 𝐶 𝑠 Comp (cid:12) (cid:12) 𝐼 𝑡,𝑠 ∈ Top-k(𝐼 𝑡,: ) . (17)
10

---

Shared Key-Value Multi-Query Attention
Concatenation
Sliding Window Heavily
KV Entries Compressed …
KV Entries
Queries
Token-Level
Compressor
…
Hidden States of KV Tokens Hidden State of Query Token
Figure4 | CorearchitecturesofHCA.Itperformsheaviercompression,wheretheKVentriesof
𝑚′ (≫ 𝑚)tokenswillbeconsolidatedintoone. Also,weadditionallyintroduceasmallsetof
slidingwindowKVentriestoenhancelocalfine-graineddependencies.
Shared Key-Value MQA. After selecting the sparse KV entries, CSA then performs core
attentioninaMulti-QueryAttention(MQA)(Shazeer,2019)manner,whereeachcompressed
KVentryin CSprsComp servesasbothattentionkeyandvalue. Tobespecific,foraquerytoken𝑡,
𝑡
𝑄
wefirstproduceattentionqueries {q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛
ℎ
} fromthecompressedlatentvectorc
𝑡
:
[q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛
ℎ
] =q𝑡 =c
𝑡
𝑄 ·𝑊𝑈𝑄 , (18)
where 𝑛 ℎ denotesthenumberofqueryheads;𝑊𝑈𝑄 ∈ R𝑑𝑐×𝑐𝑛 ℎ istheup-projectionmatricesfor
𝑄
queries. Notethatthelatentqueryvectorc issharedwiththatusedfortheindexerqueries.
𝑡
Next,weperformMQAon {q𝑡,𝑖 } and C
𝑡
SprsComp :
(cid:16) (cid:17)
o𝑡,𝑖 =CoreAttn query=q𝑡,𝑖,key=C
𝑡
SprsComp ,value=C
𝑡
SprsComp , (19)
whereo𝑡,𝑖 ∈ R𝑐 isthecoreattentionoutputofthe𝑖-thheadatthe𝑡-thtoken;CoreAttn(·) denotes
thecoreattentionoperation.
GroupedOutputProjection. IntheconfigurationofDeepSeek-V4,𝑐𝑛 ℎisquitelarge. Therefore,
directlyprojectingtheoutputsofthecoreattentionoperation [o𝑡,1 ;o𝑡,2 ;...;o𝑡,𝑛
ℎ
] =o𝑡 ∈ R𝑐𝑛 ℎ toa
𝑑-dimensionalhiddenstatewillimposeasubstantialcomputationalburden. Tomitigatethis
cost, wedesignagroupedoutputprojectionstrategy. Tobespecific, wefirstsplit 𝑛 ℎ outputs
into 𝑔 groups,andthenforeachgroupofoutputo 𝐺
𝑡,𝑖
∈ R𝑐𝑛 𝑔 ℎ ,weprojectittoa𝑑 𝑔-dimensional
intermediate output o 𝐺 𝑡,𝑖 ′ ∈ R𝑑𝑔, where 𝑑 𝑔 < 𝑐𝑛 𝑔 ℎ. Finally, we project the intermediate output
[o 𝐺
𝑡,1
′ ;o 𝐺
𝑡,2
′ ;...;o 𝐺
𝑡,𝑔
′] ∈ R𝑑𝑔𝑔 tothefinalattentionoutputoˆ𝑡 ∈ R𝑑 .
2.3.2. HeavilyCompressedAttention
The core architecture of HCA is illustrated in Figure 4, which compresses the KV cache in a
heaviermanner,butdoesnotemploysparseattention.
CompressedKey-ValueEntries. Byandlarge,thecompressionstrategyofHCAissimilarto
thatofCSA,butemploysalargercompressionrate𝑚′ (≫ 𝑚)anddoesnotperformoverlapped
11

---

compression. Let 𝐻 ∈ R𝑛×𝑑 be a sequence of input hidden states, HCA first computes the
originalKVentries𝐶 ∈ R𝑛×𝑐 andtheircorrespondingcompressionweights𝑍 ∈ R𝑛×𝑐 :
𝐶 = 𝐻·𝑊𝐾𝑉 , (20)
𝑍 = 𝐻·𝑊𝑍 , (21)
where𝑊𝐾𝑉 ,𝑊𝑍 ∈ R𝑑×𝑐 aretrainableparameters. Next,each𝑚′KVentriesin𝐶willbecompressed
into one according to the compression weights and learnable positional biases 𝐵 ∈ R𝑚′×𝑐 ,
producing𝐶Comp ∈ R 𝑚 𝑛 ′×𝑐 . Eachcompressedentry𝐶Comp ∈ R𝑐 iscomputedby
𝑖
𝑆 𝑚′𝑖:𝑚′(𝑖+1)−1 =Softmax row (𝑍 𝑚′𝑖:𝑚′(𝑖+1)−1 +𝐵), (22)
𝑚′(𝑖+1)−1
∑︁
𝐶 𝑖 Comp = 𝑆 𝑗 ⊙𝐶 𝑗. (23)
𝑗=𝑚′𝑖
Throughthiscompressionoperation,HCAcompressesthesequencelengthto 1 times.
𝑚′
SharedKey-ValueMQAandGroupedOutputProjection. HCAalsoemploysthesharedKV
MQAandgroupedoutputprojectionstrategiesasCSAdoes. AftertheKVcompression,fora
querytoken𝑡,HCAfirstproducesattentionqueries {q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛
ℎ
} inalow-rankmanner:
c
𝑡
𝑄 =h𝑡 ·𝑊𝐷𝑄 , (24)
[q𝑡,1 ;q𝑡,2 ;...;q𝑡,𝑛
ℎ
] =q𝑡 =c
𝑡
𝑄 ·𝑊𝑈𝑄 , (25)
whereh𝑡 ∈ R𝑑 istheinputhiddenstateofthequerytoken𝑡; 𝑛 ℎ denotesthenumberofquery
heads;𝑊𝐷𝑄 ∈ R𝑑×𝑑𝑐 and𝑊𝑈𝑄 ∈ R𝑑𝑐×𝑐𝑛 ℎ arethedown-projectionandup-projectionmatricesfor
queries,respectively. Next,weperformMQAon {q𝑡,𝑖 } and𝐶Comp:
(cid:16) (cid:17)
o𝑡,𝑖 =CoreAttn query=q𝑡,𝑖,key=𝐶Comp,value=𝐶Comp , (26)
whereo𝑡,𝑖 ∈ R𝑐 isthecoreattentionoutputofthe𝑖-thheadatthe𝑡-thtoken. Next,asCSAdoes,
HCAsplits𝑛 ℎ outputsinto𝑔 groups,andforeachgroupofoutputo 𝐺 𝑡,𝑖 ∈ R𝑐𝑛 𝑔 ℎ ,HCAprojectsit
toa 𝑑 𝑔-dimensionalintermediateoutputo 𝐺 𝑡,𝑖 ′ ∈ R𝑑𝑔,where 𝑑 𝑔 < 𝑐𝑛 𝑔 ℎ. Finally,HCAprojectsthe
intermediateoutput [o 𝐺
𝑡,1
′ ;o 𝐺
𝑡,2
′ ;...;o 𝐺
𝑡,𝑔
′] ∈ R𝑑𝑔𝑔 tothefinalattentionoutputoˆ𝑡 ∈ R𝑑 .
2.3.3. OtherDetails
InadditiontothecorearchitecturesofCSAandHCAdescribedabove, ourhybridattention
incorporatesseveralothertechniques. Forwritingclarity,weomittheseadditionaltechniques
fromtheaboveintroductionandwillbrieflydescribetheminthissubsection. Also,thissubsec-
tionfocusesonlyonthecoreideasofthemandmayomitsometinydetailsforsimplicity. We
encouragethereaderstorefertoouropen-sourceimplementationforunambiguousdetails.
QueryandKey-ValueEntryNormalization. ForbothCSAandHCA,weperformanaddi-
tionalRMSNormoperationoneachheadofthequeriesandtheonlyheadofthecompressedKV
entries,justbeforethecoreattentionoperation. Thisnormalizationavoidsexplodingattention
logitsandmayimprovetrainingstability.
12

---

PartialRotaryPositionalEmbedding. ForbothCSAandHCA,wepartiallyemploytheRotary
PositionalEmbedding(RoPE)(Suetal.,2024)totheattentionqueries,KVentries,andthecore
attentionoutputs. Tobespecific,foreachqueryvectorandKVentryvectorusedinCSAand
HCA, we apply RoPE to its last 64 dimensions. Since the KV entries serve as both attention
keysandvalues,thenaivecoreattentionoutputs {o𝑡,𝑖 } willcarryabsolutepositionembeddings,
derivedfromtheweightedsumofKVentries. Asacountermeasure,wealsoapplyRoPEwith
position−𝑖onthelast64dimensionsofeacho𝑡,𝑖. Inthisway,theoutputofthecoreattention
willalsocarryrelativepositionembeddings—thecontributionofeachKVentrytothecore
attentionoutputswillalsoberelatedtothedistancebetweenthequeryandtheKVentry.
AdditionalBranchofSlidingWindowAttention. Inordertostrictlypreservecausalityin
CSAandHCA,eachqueryattendstoonlyprecedingcompressedKVblocks. Consequently,a
querycannotaccessinformationfromothertokenswithinitsowncompressedblock. Meanwhile,
recenttokensusuallypossessgreaterrelevancetothequerytokeninlanguagemodeling. For
thesereasons,weintroduceasupplementaryattentionbranchtobothCSAandHCAinasliding
windowmanner,forbettermodelingoflocaldependencies. Tobespecific,foreachquerytoken,
weadditionallyproduce𝑛 uncompressedKVentriescorrespondingtotherecent𝑛 tokens.
win win
In the core attention of CSA and HCA, these KV entries in the sliding window will be used
alongwiththecompressedKVentries.
Attention Sink. In the core attention of CSA and HCA, we employ the trick of attention
sink (OpenAI, 2025; Xiao et al., 2024). To be specific, we set a series of learnable sink logits
{𝑧′,𝑧′,...,𝑧′ }. For the ℎ-th attention head, Exp(𝑧′) will be added to the denominator of the
1 2 𝑛 ℎ ℎ
attentionscore:
𝑠 ℎ,𝑖,𝑗 = (cid:205) 𝑘Exp
E
(𝑧
x
ℎ
p
,𝑖
(
,𝑘
𝑧
)
ℎ
+
,𝑖,𝑗
E
)
xp(𝑧 ℎ ′) , (27)
where 𝑠 ℎ,𝑖,𝑗,𝑧 ℎ,𝑖,𝑗 ∈ R denote the attention score and attention logit of the ℎ-th attention head
betweenthe𝑖-thquerytokenandthe 𝑗-thprecedingtokenorcompressedblock. Thistechnique
allowseachqueryheadtoadjustitstotalattentionscorestobenotequalto1,andeventobe
near0.
2.3.4. EfficiencyDiscussion
Due to the employment of hybrid CSA and HCA, together with low-precision computation
andstorage,theattentionmoduleofDeepSeek-V4seriesachievesremarkableefficiencyinboth
attention FLOPs and KV cache size, especially in long-context scenarios. First, we adopt a
mixedstorageformatforKVentries: BF16precisionisusedfortherotarypositionalembedding
(RoPE)dimensions,whileFP8precisionisappliedtotheremainingdimensions. Thishybrid
representation reduces the KV cache size by nearly half compared with pure BF16 storage.
Second, attention computation within the lightning indexer is performed in FP4 precision,
which accelerates the attention operation under extremely long contexts. Third, relative to
DeepSeek-V3.2,asmallerattentiontop-kischoseninDeepSeek-V4series,therebyimproving
modelefficiencyonshort-andmedium-lengthtexts. Finally,andmostimportantly,compressed
attentionandhybridattentiontechniquessubstantiallyreduceboththeKVcachesizeandthe
computationalFLOPs.
TakingBF16GQA8(Ainslieetal.,2023)withaheaddimensionof128asthebaseline—one
ofthecommonconfigurationsofLLMattention—theKVcachesizeofDeepSeek-V4seriescan
bedramaticallyreducedtoapproximately2%timesofthatbaselineinthe1M-contextsetting.
13

---

Algorithm1MuonOptimizerforDeepSeek-V4
Require: Learningrate𝜂,momentum 𝜇,weightdecay 𝜆,updaterescalingfactor𝛾
1: foreachtrainingstep𝑡 do
2: foreachlogicallyindependentweight𝑊 ∈ R𝑛×𝑚 do
3: 𝐺 𝑡 =∇ 𝑊 L 𝑡 (𝑊 𝑡−1 ) ⊲Computegradients
4:
𝑀
𝑡
=𝜇𝑀
𝑡−1
+𝐺
𝑡
⊲Accumulatemomentumbuffer
5: 𝑂 𝑡 ′ =HybridNewtonSchulz(𝜇𝑀 𝑡 +𝐺 𝑡 ) ⊲NesterovtrickandhybridNewton-Schulz
√︁
6:
𝑂
𝑡
=𝑂
𝑡
′· max(𝑛,𝑚)·𝛾 ⊲RescaletheupdateRMS
7:
𝑊
𝑡
=𝑊
𝑡−1
·(1−𝜂𝜆)−𝜂𝑂
𝑡
⊲Performweightdecayandupdate
8: endfor
9: endfor
Moreover,evenwhencomparedwithDeepSeek-V3.2(DeepSeek-AI,2025)—alreadyanefficient
baseline—DeepSeek-V4seriesstillexhibitssubstantialadvantagesinefficiency. Acomparison
oftheirinferenceFLOPsandKVcachesizeisprovidedintherightpartofFigure1.
2.4. MuonOptimizer
WeemploytheMuon(Jordanetal.,2024;Liuetal.,2025)optimizerforthemajorityofmodules
inDeepSeek-V4seriesduetoitsfasterconvergenceandimprovedtrainingstability. Thefull
algorithmofourMuonoptimizationissummarizedinAlgorithm1.
BasicConfigurations. WemaintaintheAdamW(LoshchilovandHutter,2017)optimizerfor
the embedding module, the prediction head module, the static biases and gating factors of
mHCmodules,andtheweightsofallRMSNormmodules. Allothermodulesareupdatedwith
Muon. FollowingLiuetal.(2025), wealsoapplyweightdecaytoMuonparameters, usethe
Nesterov(Jordanetal.,2024;Nesterov,1983)trick,andrescaletheRootMeanSquare(RMS)of
theupdatematrixforreutilizationofourAdamWhyper-parameters. Differentfromthem,we
usehybridNewton-Schulziterationsfororthogonalization.
HybridNewton-SchulzIterations. Foragivenmatrix 𝑀,letitsSingularValueDecomposition
(SVD)be 𝑀 =𝑈Σ𝑉𝑇 . TheNewton-Schulziterationsaimtoapproximatelyorthogonalize 𝑀 tobe
𝑈𝑉𝑇 . Usually, 𝑀 willbefirstnormalizedas 𝑀 0 = 𝑀/||𝑀|| 𝐹 toensureitsmaximumsingularvalue
doesnotexceed1. Then,eachNewton-Schulziterationperformsthefollowingoperation:
𝑀 𝑘 =𝑎𝑀 𝑘−1 +𝑏(𝑀 𝑘−1 𝑀 𝑘 𝑇 −1 )𝑀 𝑘−1 +𝑐(𝑀 𝑘−1 𝑀 𝑘 𝑇 −1 )2𝑀 𝑘−1 . (28)
OurhybridNewton-Schulzperforms10iterationsovertwodistinctstages. Duringthefirst8
steps,weusecoefficients (𝑎,𝑏,𝑐) = (3.4445,−4.7750,2.0315) todriverapidconvergence,bringing
thesingularvaluescloseto1. Inthefinal2steps,weswitchtocoefficients (𝑎,𝑏,𝑐) = (2,−1.5,0.5),
whichstabilizethesingularvaluespreciselyat1.
AvoidingExplodingAttentionLogits. TheattentionarchitectureofDeepSeek-V4seriesal-
lowsustodirectlyapplyRMSNormontheattentionqueriesandKVentries,whicheffectively
preventsattentionlogitsfromexploding. Consequently,wedonotemploytheQK-Cliptech-
nique(Liuetal.,2025)inourMuonoptimizer.
14

---

3. General Infrastructures
3.1. Fine-GrainedCommunication-ComputationOverlapinExpertParallelism
Mixture-of-Experts (MoE) can be accelerated via Expert Parallelism (EP). However, EP re-
quirescomplexinter-nodecommunicationandimposessubstantialdemandsoninterconnect
bandwidthandlatency. ToalleviatethecommunicationbottleneckinEPandachievehigher
end-to-end performance under lower interconnection bandwidth requirements, we propose
afine-grainedEPschemethatfusescommunicationandcomputationintoasinglepipelined
kernelforcommunication-computationoverlapping.
Communication Latency Can Be Hidden. The key insight of our EP scheme is that the
communicationlatencycanbeeffectivelyhiddenbeneathcomputationinMoElayers. Asshown
inFigure5,inDeepSeek-V4series,eachMoElayercanbedecomposedmainlyintofourstages:
twocommunication-boundstages,DispatchandCombine,andtwocomputation-boundstages,
Linear-1 and Linear-2. Our profiling reveals that within a single MoE layer, the total time of
communicationislessthanthatofthecomputation. Therefore,afterfusingcommunicationand
computationintoaunifiedpipeline,computationremainsthedominantbottleneck,implying
that the system can tolerate lower interconnect bandwidth without degrading end-to-end
performance.
(a) Naive Solution
L1 Act L2
(b) Comet
Communication
Theoretical speedup: 1.42×
Computation L1 Act L2
(c) Ours
Dispatch Dispatch All-to-All
Computation L1 L2 L1 L2 L1 L2 Theoretical speedup: 1.92× Linear 1 GEMM
SwiGLU + FP8 Cast
Activation Act Act Act Linear 2 GEMM
& Combine Combine All-to-All
Expert Wave 1 Expert Wave 2 Expert Wave 3
Figure5|IllustrationofourEPschemewithrelatedworks. Comet(Zhangetal.,2025b)overlaps
DispatchwithLinear-1,andLinear-2withCombine,separately. OurEPschemeachievesafiner-
grainedoverlappingbysplittingandschedulingexpertsintowaves. Thetheoreticalspeedupis
evaluatedintheconfigurationoftheDeepSeek-V4-Flasharchitecture.
Fine-Grained EP Scheme. To further lower the interconnect bandwidth requirement and
amplifythebenefitsofoverlapping,weintroduceafiner-grainedexpertpartitioningscheme.
Inspiredbymanyrelatedworks(Aimuyoetal.,2025;Zhangetal.,2025b),wesplitandschedule
theexpertsintowaves. Eachwaveconsistsofasmallportionofexperts. Assoonasallexperts
withinthewavehavecompletedtheircommunication,computationcancommenceimmediately
withoutwaitingforotherexperts. Insteadystate,computationofcurrentwave,tokentransferfor
thenextwave,andresultsendingofcompletedexpertsallproceedconcurrently,asdemonstrated
inFigure5. Thisformsafine-grainedpipelineamongexperts,keepingbothcomputationand
communicationcontinuousthroughoutthewave. Thewave-basedschedulingspeedsupthe
15

---

performance on extreme cases such as Reinforcement Learning (RL) rollout, which usually
encounterslong-tailsmallbatches.
Performance and Open-Sourced Mega-Kernel. We validated the fine-grained EP scheme
on both NVIDIA GPUs and HUAWEI Ascend NPUs platforms. Compared against strong
non-fused baselines, it achieves 1.50 ∼ 1.73× speedup for general inference workloads, and
upto1.96×forlatency-sensitivescenariossuchasRLrolloutsandhigh-speedagentserving.
Wehaveopen-sourcedtheCUDA-basedmega-kernelimplementationnamedMegaMoE2 asa
componentofDeepGEMM.
ObservationsandProposals. Weshareobservationsandlessonsfromkerneldevelopment
andoffersomeproposalstohardwarevendors,inthehopeofaidingefficienthardwaredesign
andachievingbettersoftware-hardwareco-design:
• Computation-CommunicationRatio. Fullcommunication-computationoverlaphinges
onthecomputation-communicationratio,ratherthanthebandwidthsolely. Denotingpeak
computethroughputas𝐶 andinterconnectbandwidthas 𝐵,communicationcanbefully
hiddenwhen𝐶/𝐵 ⩽ 𝑉 /𝑉 ,where𝑉 denotesthecomputationvolumeand𝑉
comp comm comp comm
refers to the communication volume. For DeepSeek-V4-Pro, where each token-expert
pairrequires6ℎ𝑑 FLOPs(SwiGLUgate,up,anddownprojections)butonly3ℎbytesof
communication(FP8Dispatch+BF16Combine),thissimplifiesto:
𝐶
⩽ 2𝑑 =6144 FLOPs/Byte.
𝐵
Thatis,eachGBpsofinterconnectbandwidthsufficestohidethecommunicationfor6.1
TFLOP/sofcompute. Oncebandwidthmeetsthisthreshold,itceasestobethebottleneck,
and devoting additional silicon area to further bandwidth brings diminishing returns.
We encourage future hardware designs to target such balance points rather than scale
bandwidthunconditionally.
• Power Budget. Extreme kernel fusion drives compute, memory, and network to high
loadsimultaneously,makingpowerthrottlingakeyperformancelimiter. Wesuggestthat
future hardware designs provide sufficient power headroom for such fully concurrent
workloads.
• CommunicationPrimitives. Weadoptapull-basedapproachwhereeachGPUactively
reads data from remote GPUs, avoiding the high notification latency that fine-grained
pushentails. Futurehardwarewithlower-latencycross-GPUsignalingwouldmakepush
viableandenablemorenaturalcommunicationpatterns.
• ActivationFunction. WeproposereplacingSwiGLUwithalow-costelement-wiseactiva-
tionthatinvolvesnoexponentialordivisionoperations. Thislightensthepost-GEMM
processingdirectly,andunderthesameparameterbudget,removingthegateprojection
enlargestheintermediatedimension𝑑,furtherrelaxingthebandwidthrequirement.
3.2. FlexibleandEfficientKernelDevelopmentwithTileLang
Inpractice,ourelaboratemodelarchitecturewouldhaveresultedinhundredsoffine-grained
TorchATenoperators. WeadoptTileLang(Wangetal.,2026)todevelopasetoffusedkernels
to replace the vast majority of them, delivering optimal performance with minimal effort. It
2https://github.com/deepseek-ai/DeepGEMM/pull/304
16

---

alsoallowsustoquicklyprototypeoperatorslikeattentionvariantsduringvalidation. These
kernelsplaycriticalrolesinmodelarchitecturedevelopment,large-scaletraining,andultimately
productiondeploymentofinferenceservices. AsaDomain-SpecificLanguage(DSL),TileLang
balancesdevelopmentproductivitywithruntimeefficiency,enablingrapiddevelopmentwhile
supporting deep, iterative optimizations within the same codebase. Additionally, we collab-
oratecloselywiththeTileLangcommunitytofosteramoreagile, efficient, andstablekernel
developmentworkflow.
Reducing Invocation Overhead with Host Codegen. As accelerators continue to grow in
performance, CPU-side orchestration overhead becomes increasingly prominent. For small,
highlyoptimizedkernels,suchfixedhostoverheadcaneasilycaputilizationandthroughput.
Acommonsourceofthisoverheadisthathost-sidelogic,suchasruntimecontractchecks,is
typicallywritteninPythonforflexibilityandthusincursafixedper-invocationcost.
WemitigatethisoverheadwithHostCodegen,whichmovesmosthost-sidelogicintogen-
erated host code. Specifically, we first co-generate the device kernel and a lightweight host
launcherattheIR(IntermediateRepresentation)level,embeddingthenecessarymetadata—such
asdatatypes,rank/shapeconstraints,andstride/layoutassumptions—parsedfromthelan-
guage frontend. The launcher is then lowered to the host source code built on top of the
TVM-FFI(Chenetal.,2018)framework,whosecompactcallingconventionandzero-copytensor
interoptogetherminimizehost-sideoverhead. Atruntime,thisgeneratedhostcodeperforms
validationandargumentmarshaling,shiftingallper-invocationchecksoutofthePythonexe-
cutionpath. OurmeasurementsshowthatCPU-sidevalidationoverheaddropsfromtensor
hundredsofmicrosecondstolessthanonemicrosecondperinvocation.
SMT-Solver-Assisted Formal Integer Analysis. TileLang kernels involve complex tensor
indexarithmeticthatrequiresstrongformalintegeranalysis. Duringcompilationpassessuch
aslayoutinference,memoryhazarddetection,andboundanalysis,thecompilermustverify
whetherintegerexpressionssatisfyspecificpropertiestoenablethecorrespondingoptimiza-
tions. Therefore,strongerformalanalysiscapabilitiescanunlockmoreadvancedandcomplex
optimizationopportunities.
Tothisend,weintegratetheZ3SMTsolver(DeMouraandBjørner,2008)intoTileLang’s
algebraicsystem,providingformalanalysiscapabilityformostintegerexpressionsintensor
programs. Westrikeabalancebetweencomputationaloverheadandformalexpressivenessby
translatingTileLang’sintegerexpressionsintoZ3’squantifier-freenon-linearintegerarithmetic
(QF_NIA).BasedonIntegerLinearProgramming(ILP)solvers,QF_NIAseamlesslyresolves
standardlinearintegerexpressionscommoninkernels. Furthermore,itsinherentnon-linear
reasoningcapacityeffectivelyaddressesadvancedchallengeslikevectorizationovervariable
tensorshapes. Underreasonableresourcelimits,Z3elevatesoveralloptimizationperformance
while restricting compilation time overhead to just a few seconds. The impact is substantial
acrossmultiplepasses,includingvectorization,barrierinsertion,andcodesimplification.
NumericalPrecisionandBitwiseReproducibility. Inproductionsettings,numericalcorrect-
nessandreproducibilityareascriticalasrawthroughput. Wethereforeprioritizeaccuracyby
default: fast-mathoptimizationsaredisabledatthecompilerlevel,andprecision-affectingap-
proximationsareprovidedonlyasexplicit,opt-infrontendoperators(e.g.,T.__exp,T.__log,
and T.__sin). Conversely, when strict IEEE-754 semantics are required, TileLang provides
17

---

IEEE-compliantintrinsicswithexplicitroundingmodes(e.g.,T.ieee_fsqrt,T.ieee_fdiv,
andT.ieee_add),enablingdeveloperstopreciselyspecifynumericalbehavior.
Wealsotargetbitwisereproducibilityforvalidatingkernelsagainsthand-writtenCUDA
baselines. We align TileLang’s algebraic simplification and lowering rules with mainstream
CUDAtoolchains(e.g.,NVCC)toavoidtransformationsthatintroduceunintendedbit-level
differences. Layoutannotations(e.g.,T.annotate_layout)furtherallowuserstopindown
layout-dependentloweringdecisions,keepingevaluationandaccumulationorderconsistent
withthereferenceCUDAimplementationandthusenablingbit-identicaloutputswhendesired.
Ourevaluationshowsthattheseaccuracy-andreproducibility-orienteddesignchoicesdo
notsacrificeperformance: underconservativedefaults,TileLangkernelsremaincompetitive,
whileexposingknobstoselectivelyrelaxnumericalconstraintsforhigherspeed.
3.3. High-PerformanceBatch-InvariantandDeterministicKernelLibraries
Toenableefficienttrainingandinference,wedevelopacomprehensivesetofhigh-performance
computational kernels. Beyond basic functionalities and maximizing hardware utilization,
anotherpivotaldesigngoalistoensuretrainingreproducibilityandbitwisealignmentamong
pre-training, post-training, and inference pipelines. Therefore, we implement end-to-end,
bitwisebatch-invariant,anddeterministickernelswithminimalperformanceoverhead. These
kernelsarehelpfulfordebugging,stabilityanalysis,andconsistentpost-trainingbehavior.
BatchInvariance. Batchinvarianceensuresthattheoutputofanygiventokenremainsbitwise
identical,regardlessofitspositionwithinabatch. Toimplementbatchinvariance,theprimary
challengesarelistedasfollows:
• Attention. Toachievebatchinvariance, wecannotusethesplit-KVmethod(Daoetal.,
2023),whichdistributestheattentioncomputationforasinglesequenceacrossmultiple
Stream Multiprocessors (SMs) to balance the load of SMs. However, abandoning this
techniquewillleadtoseverewave-quantizationproblems3,whichcanadverselyaffect
GPUutilization. Toaddressthis,wedevelopadual-kernelstrategyforbatch-invariant
decoding. Thefirstkernelcomputestheattentionoutputforanentiresequencewithin
asingleSM,ensuringhighthroughputforfullyoccupiedwaves. Thesecondkernel,to
minimizethelatencyofthefinalpartially-filledwaveandthusalleviatewave-quantization,
uses multiple SMs for a single sequence. For the bitwise identity of these two kernels,
wecarefullydesignthecalculationpathofthesecondkerneltoensureitsaccumulation
orderisthesameasthatofthefirstkernel. Additionally,thesecondkernelutilizesdis-
tributedsharedmemory4 withinthread-blockclusters,enablinghigh-speeddataexchange
acrossSMs. Thisdual-kernelmethodeffectivelyconfinestheoverheadofbatch-invariant
decodingtobenegligible.
• MatrixMultiplication. TraditionalcuBLASlibrary(NVIDIACorporation,2024)cannot
achieve batch invariance. Therefore, we replace it end-to-end with DeepGEMM (Zhao
etal.,2025). Furthermore,forverysmallbatchsizes,conventionalimplementationusually
employssplit-k(Osamaetal.,2023)techniquestoimproveperformance. Unfortunately,
split-ktechniquescannotguaranteebatchinvariance,apivotalfeatureinDeepSeek-V4.
3https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplicat
ion/index.html#wave-quant
4https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/writing-cuda-kernels
.html#distributed-shared-memory
18

---

Therefore,weabandonsplit-kinmostscenarios,which,however,maycauseperformance
degradation. Toaddressthis,weintroduceasetofoptimizationsthatenableourimple-
mentationofmatrixmultiplicationtomatchorevensurpasstheperformanceofstandard
split-kinmostmajorscenarios.
Determinism. Deterministictrainingishighlybeneficialfordebugginghardwareorsoftware
issues. Moreover,whentrainingexhibitsanomaliessuchaslossspikes,determinismenables
researcherstomoreeasilypinpointnumericalcausesandfurtherrefinethemodeldesign. Non-
determinismintrainingtypicallystemsfromnon-deterministicaccumulationorder,oftendue
totheuseofatomicadditioninstructions. Thisissueprimarilyoccursduringthebackwardpass,
notablyatthefollowingparts:
• Attention Backward. In conventional implementations of backward propagation for
sparse attention, we use atomicAdd to accumulate gradients for the KV tokens. This
introducesnon-determinismduetothenon-associativityoffloating-pointaddition. To
addressthisproblem,weallocateseparateaccumulationbuffersforeachSM,followedby
aglobaldeterministicsummationacrossallbuffers.
• MoE Backward. When multiple SMs from different ranks concurrently write data to
the same buffer on a receiving rank, negotiating writing positions also introduces non-
determinism. Toresolvethis,wedesignatokenorderpre-processingmechanismwithin
each single rank, combined with buffer isolation across multiple ranks. This strategy
ensuresdeterminismofboththesendresultsofexpertparallelismandtheaccumulation
orderintheMoEbackwardpass.
• MatrixMultiplicationinmHC.mHCinvolvesamatrixmultiplicationwithanoutputdi-
mensionofonly24. Forverysmallbatchsizes,wearecompelledtousethesplit-k(Osama
et al., 2023) algorithm, whose naive implementation will cause non-determinism. To
overcomethis,weoutputeachsplitpartseparatelyandperformadeterministicreduction
inasubsequentkernel,therebypreservingbothperformanceanddeterminism.
3.4. FP4Quantization-AwareTraining
Toachieveinferenceaccelerationandmemorysavingsatdeployment,weintroduceQuantization-
Aware Training (QAT) (Jacob et al., 2018) during the post-training stage, enabling the model
to adapt to the precision degradation introduced by quantization. We apply FP4 (MXFP4)
quantization (Rouhani et al., 2023) to two components: (1) MoE expert weights, which are a
major source of GPU memory occupancy (OpenAI, 2025), and (2) the Query-Key (QK) path
in the indexer of CSA, where QK activations are cached, loaded, and multiplied entirely in
FP4,acceleratingattentionscorecomputationinlong-contextscenarios. Inaddition,wefurther
quantize the index scores 𝐼 from FP32 to BF16 during this QAT process. This optimization
:,:
achievesa2×speedupforthetop-kselector,whilepreservinga99.7%recallrateofKVentries.
ForMoEexpertweights,followingthecommonpracticeofQAT,theFP32masterweights
maintained by the optimizer are first quantized to FP4, then dequantized back to FP8 for
computation. Notably,ourFP4-to-FP8dequantizationislossless. ThisisbecauseFP8(E4M3)
has 2 additional exponent bits compared with FP4 (E2M1), offering a larger dynamic range.
Consequently,aslongastheratiobetweenthemaximumandminimumscalefactorsoftheFP4
sub-blocks (1×32 tiles) within each FP8 quantization block (128×128 tiles) does not exceed
acertainthreshold,thefine-grainedscaleinformationcanbefullyabsorbedbytheextended
dynamicrangeofFP8. Weempiricallyverifythatcurrentweightssatisfythiscondition. This
allows the entire QAT pipeline to fully reuse the existing FP8 training framework without
19

---

anymodification. Inthebackwardpass,gradientsarecomputedwithrespecttothesameFP8
weightsintheforwardpassanddirectlypropagatedbacktotheFP32masterweights,equivalent
toapplyingtheStraight-ThroughEstimator(STE)throughthequantizationoperation. Thisalso
avoidstheneedtore-quantizetransposedweights.
During the inference and rollout phases of RL training, which do not involve backward
passes, we directly use real FP4 quantized weights instead of simulated quantization. This
ensuresthatmodelbehaviorduringsamplingisfullyconsistentwithonlinedeployment,while
alsoreducingkernelmemoryloadingforactualspeedupandsignificantlyloweringmemory
consumption. WeprocesstheQKpathintheindexerofCSAsimilarly.
3.5. TrainingFramework
Our training framework is built upon the scalable and efficient infrastructure developed for
DeepSeek-V3(DeepSeek-AI,2024). IntrainingDeepSeek-V4,weinheritthisrobustfoundation
whileintroducingseveralkeyinnovationstoaccommodateitsnovelarchitecturalcomponents—
specificallytheMuonoptimizer,mHC,andthehybridattentionmechanism—whilemaintaining
hightrainingefficiencyandstability.
3.5.1. EfficientImplementationofMuon
TheMuonoptimizerrequiresthefullgradientmatrixtocomputeparameterupdates,which
presentsachallengewhencombinedwiththeZeroRedundancyOptimizer(ZeRO)(Rajbhandari
etal.,2020). TraditionalZeROisdesignedforelement-wiseoptimizerslikeAdamW,wherea
singleparametermatrixcanbepartitionedandupdatedacrossmultipleranks. Toaddressthis
conflict,wedesignahybridstrategyofZeRObucketassignmentforMuon.
For dense parameters, we limit the maximum size of ZeRO parallelism and employ a
knapsackalgorithmtoassignparametermatricestotheseranks,ensuringeachrankmanagesa
roughlybalancedload. Thebucketoneachrankispaddedtomatchthesizeofthelargestbucket
acrossranks,facilitatingefficientreduce-scatteroperations. Thispaddingtypicallyincursless
than10%memoryoverheadinoursetup,whereeachrankmanagesnomorethanfiveparameter
matrices. When the overall size of data parallelism exceeds the limit for ZeRO, we compute
theMuonupdateredundantlyacrosstheextradata-parallelgroups,tradingcomputationfor
reducedtotalbucketmemory.
For MoE parameters, we optimize each expert independently. We first flatten all down
projection matrices in SwiGLU (Shazeer, 2020) of all experts across all layers, followed by
flattenedupprojectionmatricesandgatematrices. Then,wepadtheflattenedvectortoensure
wecanevenlydistributethisvectoracrossallrankswithoutsplittinganylogicallyindependent
matrix. Giventhelargenumberofexperts,wedonotimposealimitofZeROparallelismfor
MoEparameters,andthepaddingoverheadisalsonegligible.
Additionally,oneachrank,consecutiveparametersofidenticalshapewillbeautomatically
merged,enablingbatchedexecutionoftheNewton-Schulziterationsforbetterhardwareutiliza-
tion. Furthermore,weobservethattheNewton-SchulziterationsinMuonremainstablewhen
computedwithBF16matrixmultiplications. Leveragingthis,wefurtherquantize,inastochastic
roundingmanner,theMoEgradientstobesynchronizedacrossdata-parallelrankstotheBF16
precision, halving the communication volume. To avoid accumulation errors introduced by
low-precisionadders,wereplaceconventionaltree-orring-basedreduce-scattercollectiveswith
atwo-phaseapproach. First,anall-to-alloperationexchangeslocalgradientsacrossranks,and
theneachrankperformsalocalsuminFP32. Thisdesignmaintainsnumericalrobustness.
20

---

3.5.2. Cost-EffectiveandMemory-EfficientImplementationofmHC
TheintroductionofmHCincreasesbothactivationmemoryconsumptionandcommunication
volumebetweenpipelinestages,comparedwithconventionalresidualconnections. Tomitigate
thesecosts,weimplementseveraloptimizationstrategies.
Firstly, we carefully design and implement fused kernels of mHC for both training and
inference. Secondly,weintroducearecomputationstrategythatselectivelycheckpointsinterme-
diatetensors. Specifically,werecomputemosthiddenstatesbetweenlayersandallnormalized
layerinputs,whileavoidingrecomputationofcompute-intensiveoperations. Thisachievesa
balancebetweenmemorysavingandcomputationaloverhead. Thirdly,weadjusttheDualPipe
1F1Boverlappingschemetoaccommodatetheincreasedpipelinecommunicationandenable
concurrentexecutionofsomeoperationsinmHC.
Collectively,theseoptimizationsconstrainthewall-timeoverheadofmHCtoonly6.7%of
theoverlapped1F1Bpipelinestage. Moredetailsoftheengineeringoptimizationcanbefound
inthededicatedmHCpaper(Xieetal.,2026).
3.5.3. ContextualParallelismforLong-ContextAttention
Conventional Context Parallelism (CP) partitions the sequence dimension, with each rank
maintainingcontiguous𝑠tokens. Thisintroducestwochallengestoourcompressedattention
mechanisms(i.e.,CSAandHCA).Ontheonehand,trainingsamplesarepackedfrommultiple
sequences,andeachsequenceiscompressedindependentlybyafactorof𝑚(or𝑚′),withany
trailingtokensfewerthan 𝑚 beingdiscarded. Consequently, thecompressedKVlengthsare
typically less than 𝑠 and vary across ranks. On the other hand, the compression requires 𝑚
𝑚
consecutiveKVentries,whichmaystraddletheboundarybetweentwoneighboringCPranks.
Toaddressthesechallenges,wedesignatwo-stagecommunicationapproach. Inthefirst
stage, each rank 𝑖 sends its last 𝑚 uncompressed KV entries to rank 𝑖+1. Then, rank 𝑖+1
compressessomeofthesereceivedentriestogetherwithitslocal 𝑠uncompressedKVentries,
producingafixedlengthof 𝑠 +1compressedentries,inwhichexistsomepaddingentries. In
𝑚
thesecondstage,anall-gatheroperationacrossallCPrankscollectsthelocallycompressedKV
entries. Then,afusedselect-and-padoperatorreorganizesthemintothefullsetofcompressed
KVentrieswithatotallengthofcp_size· 𝑠 . Anypaddingentriesareplacedatthetail. For
𝑚
HCAandtheindexerinCSA,thevisiblerangeofcompressedKVentriesforeachquerytoken
can be precomputed by rules. For the sparse attention in CSA, the top-𝑘 selector explicitly
specifiestheindicesofvisiblecompressedKVentriesforeachquery.
3.5.4. ExtendedAutomaticDifferentiationforFlexibleActivationCheckpointing
Conventionalactivationcheckpointingimplementationsoperateatthegranularityofanentire
module,decidingwhethertoretainorrecomputeitsoutputactivationsduringthebackward
pass. Thiscoarsegranularityoftenleadstosuboptimaltrade-offsbetweenrecomputationcost
andactivationmemoryfootprint. Analternativeapproachistomanuallyimplementtheforward
andbackwardlogicofanentirelayer,explicitlymanagingtensorcheckpointingstates. While
enablingfine-grainedcontrol,thismethodlosestheconvenienceoftheautomaticdifferentiation
framework,substantiallyincreasingdevelopmentcomplexity.
Toachievefine-grainedcontrolwithoutsacrificingprogrammingefficiency,weimplementa
tensor-levelactivationcheckpointingmechanismwithautomaticdifferentiationsupport. With
thismechanism,developersonlyneedtoimplementtheforwardpassandselectivelyannotate
21

---

individualtensorsforautomaticcheckpointingandrecomputation. Ourframeworkleverages
TorchFX(Reedetal.,2022)totracethefullcomputationgraph. Foreachannotatedtensor, it
performsabackwardtraversaltoidentifytheminimalsubgraphrequiredforitsrecomputation.
Wedefinetheseminimalsubgraphsasrecomputationgraphsandinsertthemintothebackward
logicjustbeforethecorrespondinggradientcomputation.
Comparedwiththemanualimplementation,thisdesignintroducesnoadditionaloverhead
during training. Recomputation in this framework is implemented by directly freeing the
GPU memory of the annotated tensor and reusing the storage pointer from the recomputed
tensor,withoutanyGPUmemorycopy. Furthermore,sincegraphtracingexecutesthemodel
concretely,wecantracktheunderlyingstoragepointerofeachtensor,whichenablesautomatic
deduplicationofrecomputationfortensorsthatsharestorage(e.g.,theinputandoutputofa
reshapeoperation). Thisrelievesdevelopersfromreasoningaboutlow-levelmemorydetails
whenannotatingrecomputation.
3.6. InferenceFramework
OurinferenceframeworklargelyinheritsfromthatofDeepSeek-V3,withsomedifferencesin
KVCachemanagement.
3.6.1. KVCacheStructureandManagement
ToefficientlymanagetheheterogeneousKVcachesarisingfromthehybridattentionmechanism
inDeepSeek-V4,wedesignacustomizedKVcachelayout. ThelayoutisillustratedinFigure6,
andwewillelaborateonitindetailasfollows.
HeterogeneousKVEntriesinDeepSeek-V4. ThehybridattentionmechanisminDeepSeek-
V4 series introduces multiple types of KV entries with different Key-Value (KV) cache sizes
andupdaterules. Thelightningindexerforsparseselectionintroducesadditionaldimensions
into the KV cache that possess embedding sizes distinct from those in the primary attention.
ThecompressiontechniquesemployedinCSAandHCAreducethesequencelengthbyfactors
of 1 and 1 ,respectively,therebydecreasingtheoverallKVcachesize. Asaresult,KVcache
𝑚 𝑚′
sizesvaryacrossdifferentlayers. Furthermore,SlidingWindowAttention(SWA)layersalso
operate with distinct KV cache sizes, as well as separate cache hit and eviction policies. In
the compression branch, one KV entry is generated for every 𝑚 tokens. When the number
of remaining tokens is insufficient for compression, all pending tokens and their associated
hidden states must be retained in a buffer until the compression operation can be executed.
Thesebufferedtokensrepresentasequencestatedeterminedbypositionalcontextandarealso
managedwithintheKVcacheframework.
Challenges in Managing Hybrid Attention KV Cache. The hybrid attention mechanism
violatesfundamentalassumptionsbehindPagedAttentionanditsvariants. Althoughrecent
hybridKVcachemanagingalgorithms(e.g.,Jenga(Zhangetal.,2025a),Hymba(Dongetal.,
2025)) target general hybrid attention models or specific structures, two principal obstacles
preventconsolidatingKVcachesacrossalllayersunderthePagedAttentionframework:
• Diversecachepolicies,suchasthoseusedinSlidingWindowAttention.
• Constraintsimposedbyhigh-performanceattentionkernels,includingalignmentrequire-
ments.
22

---

State Cache KV Cache
Uncompressed
Request 1 SWA KV Block 0
KV State
Uncompressed
Request 2 SWA KV Block 1 KV State
Uncompressed
Request 3 SWAL aKyVer-0 SWA KV Layer-2 CSA State Block 2
KV State
…UncompresseLdayer-3 HCA State
SWA KV KV State
Layer-n SWA KV …
Uncompressed
Request R SWA KV Block N
KV State
……
CSA KV
HCA KV
LaCySeAr- 2KV CSA Indexer KV CSA Main KV
HCA KV of k1 tokens of k1 tokens
La C y S e A r- 3 KV HCA KV of k2 tokens
LaHyCeAr- 4KV CSA Indexer KV CSA Main KV
CSA KV of k1 tokens of k1 tokens
La H y C e A r- 5 KV HCA KV of k2 tokens
... ...
CSA KV
HCA KV
……
Figure6 | IllustrationoftheKVcacheLayoutforDeepSeek-V4. TheKVcacheisorganizedinto
twoprimarycomponents: aclassicalKVcacheforCSA/HCA,andastatecacheforSWAand
unready-for-compressiontokensinCSA/HCA.Inthestatecache,eachrequestisassigneda
fixed-sizecacheblock. Withinthisblock,theSWAsegmentstorestheKVentriescorresponding
tothemostrecent𝑛 tokens,whiletheCSA/HCAsegmentstoresuncompressedtailstates
win
thatarenotyetreadyforcompression. IntheclassicalKVcache,weallocatemultipleblocks
perrequest. Eachcacheblockcoverslcm(𝑚,𝑚′) originaltokens,producing𝑘 = lcm(𝑚,𝑚′) CSA
1 𝑚
compressedtokensand𝑘 =
lcm(𝑚,𝑚′)
HCAcompressedtokens.
2 𝑚′
For efficient KV cache management of DeepSeek-V4, we design corresponding strategies to
overcomethesetwochallenges.
StateCacheforSWAandUncompressedTailTokens. Toaddressthefirstobstacle,weadopt
analternativecachemanagementmechanism. SinceSWAisdesignedtoenhanceperformance
underalimitedKVcachesize,itisreasonabletotreatit,alongwiththeuncompressedtailtokens
fromthecompressionbranch,asastate-spacemodel. ThecorrespondingKVcachecanthusbe
regardedasasequence-specificstatethatdependssolelyonthecurrentposition. Accordingly,
wepre-allocateafixed-andlimited-sizepoolofstatecaches,anddynamicallyassignittoeach
sequence.
Sparse Attention Kernel Co-Design. Regarding the second obstacle, conventional high-
performanceattentionkernelstypicallyassumeafixednumber 𝐵oftokensperblocktooptimize
performance, correspondingto 𝐵·𝑚 originaltokensinCSAand 𝐵·𝑚′ inHCA.Throughem-
ployingahigh-performancesparse-attentionkernel,differentlayerscanaccommodatevariable
tokensperblockwithoutperformancedegradation. Achievingthisrequiresco-designingthe
KV cache layout and the sparse attention kernel. For instance, padding blocks to align with
cachelinescanimproveperformance. Thus,forCSAwithcompressionratio𝑚andHCAwith
ratio 𝑚′, thenumberoforiginaltokensperblockcanbeanymultipleoflcm(𝑚,𝑚′), theleast
commonmultipleofthesetwocompressionratios.
3.6.2. On-DiskKVCacheStorage
WhenservingDeepSeek-V4,weleverageanon-diskKVcachestoragemechanismtoeliminate
repeatedprefillingforshared-prefixrequests. ForthecompressedKVentriesinCSA/HCAand
theuncompressedKVentriesinSlidingWindowAttention(SWA),wedesignseparatesolutions
forstoragemanagement.
23

---

For CSA and HCA, we simply store all of the compressed KV entries to the disk. When
a request hits a stored prefix, we read and reuse the compressed KV entries corresponding
totheprefix,untilthelastcompletecompressionblock. Specially,forprefixtokensinthetail
incompleteblock,westillneedtorecomputethemtorestoretheuncompressedKVentries,as
uncompressedKVentriesinCSAandHCAarenotstored.
FortheSWAKVentries,sincetheyarenotcompressedandexistineverylayer,theirvolume
is approximately 8 times larger than the compressed CSA and HCA KV entries. To handle
theselargeSWAKVentriesefficiently,weproposeandimplementthreedistinctstrategiesfor
managingon-diskSWAKVentries,eachofferingadifferenttrade-offbetweenstorageoverhead
andcomputationalredundancy:
• Full SWA Caching. This strategy stores the complete SWA KV entries for all tokens,
ensuringcomputationalzero-redundancy. Underthisstrategy,theSWAKVentriesofthe
hittingprefixcanbereconstructedbyjustreadingtheon-diskcacheofthelast𝑛 tokens
win
withinthatprefix. Despitecomputationalzero-redundancy,thisstrategyisinefficientfor
modernSSD-basedstoragesystems—onlyasmallsubsetofthestoredSWAKVcache
will be accessed for each hitting request, which leads to an unbalanced write-intensive
accesspattern.
• PeriodicCheckpointing. ThisstrategycheckpointsSWAKVentriesofthelast𝑛 tokens
win
withinevery 𝑝tokens,where 𝑝isatunableparameter. Forahittingprefix,weloadthe
mostrecentcheckpointedstate,andthenrecomputetheremainingtailtokens. Through
tuning 𝑝,thisstrategyenablesanon-demandtrade-offbetweenstorageandcomputation.
• ZeroSWACaching. ThisstrategydoesnotstoreanySWAKVentries. Forahittingprefix,
weneedtoperformmorerecomputationtorestoretheSWAKVentries. Tobespecific,in
eachattentionlayer,theSWAKVentryofeachtokendependsontheSWAKVentriesof
onlythemostrecent 𝑛 tokensfromthepreviouslayer. Therefore, leveragingcached
win
CSAandHCAKVentries,recomputingthelast𝑛 ·𝐿tokensisenoughtorestorethelast
win
𝑛 SWAKVentriesforan 𝐿-layermodel.
win
Dependingonspecificdeploymentscenarios,weselectthemostsuitablestrategytoachievethe
desiredtrade-offbetweenstorageandcomputation.
4. Pre-Training
4.1. DataConstruction
Ontopofthepre-trainingdataofDeepSeek-V3,weendeavortoconstructamorediverseand
higher-qualitytrainingcorpuswithlongereffectivecontexts. Wecontinuallyrefineourdatacon-
structionpipelines. Forweb-sourceddata,weimplementfilteringstrategiestoremovebatched
auto-generatedandtemplatedcontent,therebymitigatingtheriskofmodelcollapse(Zhuetal.,
2024). Mathematicalandprogrammingcorporastillremaincorecomponentsofourtraining
data,andwefurtherenhancethecodingcapabilitiesofDeepSeek-V4seriesbyincorporating
agentic data during the mid-training phase. For multilingual data, we build a larger corpus
forDeepSeek-V4,improvingitscaptureoflong-tailknowledgeacrossdifferentcultures. For
DeepSeek-V4, we place a particular emphasis on long-document data curation, prioritizing
scientific papers, technical reports, and other materials that reflect unique academic values.
Combiningalltheabove,ourpre-trainingcorpuscomprisesmorethan32Ttokens,containing
mathematicalcontents,codes,webpages,longdocuments,andotherhigh-qualitycategories.
For pre-training data, we largely follow the same pre-processing strategies of DeepSeek-
24

---

V3. Fortokenization,ontopoftheDeepSeek-V3tokenizer,weintroduceafewspecialtokens
for context construction, and still remain the vocabulary size to be 128K. We also inherit the
token-splitting (DeepSeek-AI, 2024) and Fill-in-Middle (FIM) (DeepSeek-AI, 2024) strategies
fromDeepSeek-V3. InspiredbyDingetal.(2024),wepackdocumentsfromdifferentsources
intoappropriatesequencestominimizesampletruncation. DifferentfromDeepSeek-V3,we
employsample-levelattentionmaskingduringpre-training.
4.2. Pre-TrainingSetups
4.2.1. ModelSetups
DeepSeek-V4-Flash. WesetthenumberofTransformerlayersto43andthehiddendimension
𝑑 to4096. Forthefirsttwolayers,weusepureslidingwindowattention. Forthesubsequent
layers,CSAandHCAareusedinaninterleavedmanner. ForCSA,wesetthecompressionrate
𝑚to4,thenumberofindexerqueryheads𝑛𝐼 to64,theindexerheaddimension𝑐𝐼
to128,and
ℎ
thenumberofKVentriesselectedforsparseattention(i.e.,attentiontop-k)to512. ForHCA,
we set the compression rate 𝑚′ to 128. For both CSA and HCA, we set the number of query
heads𝑛 ℎ to64,theheaddimension𝑐to512,andthequerycompressiondimension𝑑 𝑐 to1024.
Thenumberofoutputprojectiongroups𝑔 issetto8,andthedimensionofeachintermediate
attentionoutput 𝑑 𝑔 issetto1024. Fortheadditionalbranchofslidingwindowattention, the
window size 𝑛 is set to 128. We employ MoE layers in all Transformer blocks, but use the
win
Hashroutingstrategyforthefirst3MoElayers. EachMoElayerconsistsof1sharedexpertand
256routedexperts,wheretheintermediatehiddendimensionofeachexpertis2048. Amongthe
routedexperts,6expertswillbeactivatedforeachtoken. Themulti-tokenpredictiondepthis
setto1. AsformHC,theexpansionfactor𝑛 issetto4,andthenumberofSinkhorn-Knopp
hc
iterations𝑡 issetto20. Underthisconfiguration,DeepSeek-V4-Flashcomprises284Btotal
max
parameters,ofwhich13Bareactivatedforeachtoken.
DeepSeek-V4-Pro. WesetthenumberofTransformerlayersto61andthehiddendimension
𝑑 to7168. Forthefirsttwolayers,weuseHCA.Forthesubsequentlayers,CSAandHCAare
used in an interleaved manner. For CSA, we set the compression rate 𝑚 to 4, the number of
indexerqueryheads𝑛𝐼 to64,theindexerheaddimension𝑐𝐼
to128,andthenumberofKVentries
ℎ
selected for sparse attention (i.e., attention top-k) to 1024. For HCA, we set the compression
rate𝑚′ to128. ForbothCSAandHCA,wesetthenumberofqueryheads𝑛 ℎ to128,thehead
dimension 𝑐 to512, andthequerycompressiondimension 𝑑 𝑐 to1536. Thenumberofoutput
projectiongroups𝑔issetto16,andthedimensionofeachintermediateattentionoutput𝑑 𝑔 isset
to1024. Fortheadditionalbranchofslidingwindowattention,thewindowsize𝑛 issetto
win
128. WeemployMoElayersinallTransformerblocks,butusetheHashroutingstrategyforthe
first3MoElayers. EachMoElayerconsistsof1sharedexpertand384routedexperts,where
theintermediatehiddendimensionofeachexpertis3072. Amongtheroutedexperts,6experts
willbeactivatedforeachtoken. Themulti-tokenpredictiondepthissetto1. AsformHC,the
expansionfactor𝑛 issetto4,andthenumberofSinkhorn-Knoppiterations𝑡 issetto20.
hc max
Underthisconfiguration,DeepSeek-V4-Procomprises1.6Ttotalparameters,ofwhich49Bare
activatedforeachtoken.
4.2.2. TrainingSetups
DeepSeek-V4-Flash. WeemploytheMuonoptimizer(Jordanetal.,2024;Liuetal.,2025)for
themajorityofparameters,butusetheAdamWoptimizer(LoshchilovandHutter,2017)forthe
25

---

embeddingmodule,thepredictionheadmodule,andtheweightsofallRMSNormmodules. For
AdamW,wesetitshyper-parametersto 𝛽 =0.9, 𝛽 =0.95,𝜀 =10−20,andweight_decay =0.1.
1 2
ForMuon,wesetthemomentumto0.95andtheweightdecayto0.1,andrescaletheRMSofeach
updatematrixto0.18forreutilizationoftheAdamWlearningrate. WetrainDeepSeek-V4-Flash
on32Ttokens, andasinDeepSeek-V3, wealsoemployabatchsizeschedulingstrategythat
increasesthebatchsize(intokens)fromasmallsizeto75.5Mandthenkeepsitat75.5Mduring
mostofthetraining. Thelearningrateislinearlywarmedupinthefirst2000steps,maintained
at2.7×10−4 formostofthetraining. Neartheendofthetraining,wefinallydecaythelearning
rateto2.7×10−5 followingacosineschedule. Thetrainingstartswithasequencelengthof4K,
andwegraduallyextendthetrainingsequencelengthto16K,64K,and1M.Asforthesetupsof
sparseattention,wefirstwarmupthemodelwithdenseattentionforthefirst1Ttokens,and
introducesparseattentionatthesequencelengthof64Kandkeepsparseattentionduringthe
restofthetraining. Whenintroducingattentionsparsity,wefirstsetashortstagetowarmup
the lightning indexer in CSA, and then train the model with sparse attention for most of the
training. Forauxiliary-loss-freeloadbalancing,wesetthebiasupdatespeedto0.001. Forthe
balanceloss,wesetitslossweightto0.0001toavoidextremeimbalancewithinsinglesequences.
TheMTPlossweightissetto0.3formostofthetraining,andto0.1uponthestartoflearning
ratedecay.
DeepSeek-V4-Pro. Except for specific values of hyper-parameters, the training setup of
DeepSeek-V4-ProislargelyconsistentwiththatofDeepSeek-V4-Flash. WeemploytheMuonop-
timizerforthemajorityofparameters,butusetheAdamWoptimizerfortheembeddingmodule,
thepredictionheadmodule,andtheweightsofallRMSNormmodules. Thehyper-parameters
ofAdamWandMuonarethesameasthoseofDeepSeek-V4-Flash. WetrainDeepSeek-V4-Pro
on 33T tokens, and also employ a batch size scheduling strategy, with the maximum batch
size being 94.4M tokens. The learning rate scheduling strategy is largely the same as that of
DeepSeek-V4-Flash,butthepeaklearningrateissetto2.0×10−4 andtheendlearningrateisset
to2.0×10−5. Thetrainingalsostartswithasequencelengthof4K,andthelengthisgradually
extendedto16K,64K,and1M.ComparedwithDeepSeek-V4-Flash, DeepSeek-V4-Prostarts
withalongerstageofdenseattention,andthestrategyofintroducingsparseattentionisthe
sameasDeepSeek-V4-Flash,followingatwo-stagetrainingmethod. Forauxiliary-loss-freeload
balancing,wesetthebiasupdatespeedto0.001. Forthebalanceloss,wesetitslossweightto
0.0001toavoidextremeimbalancewithinsinglesequences. TheMTPlossweightissetto0.3for
mostofthetraining,andto0.1uponthestartoflearningratedecay.
4.2.3. MitigatingTrainingInstability
Trainingtrillion-parameterMoEmodelspresentssignificantstabilitychallenges,andDeepSeek-
V4 series are no exception. We encountered notable instability challenges during training.
Whilesimplerollbackscouldtemporarilyrestorethetrainingstate,theyprovedinadequateasa
long-termsolutionbecausetheydonotpreventtherecurrenceoflossspikes. Empirically,we
identifiedthattheoccurrenceofspikesisconsistentlytiedtooutliersintheMoElayers,andthe
routingmechanismitselfappearstoexacerbatetheemergenceoftheseoutliers. Therefore,we
soughttotacklethisissuefromtwodimensions: breakingtheviciouscycleinducedbyrouting,
anddirectlysuppressinganomalousvalues. Fortunately,wediscoveredtwopracticaltechniques
thateffectivelymaintaintrainingstability. Althoughacomprehensivetheoreticalunderstanding
oftheirunderlyingmechanismsremainsanopenquestionfornow,wearesharingthemopenly
tofosterfurtherexplorationbythecommunity.
26

---

AnticipatoryRouting. Wefoundthatdecouplingthesynchronousupdatesofthebackbone
networkandtheroutingnetworksignificantlyimprovestrainingstability. Consequently,atstep
𝑡,weusethecurrentnetworkparameters𝜃 𝑡 forfeaturecomputation,buttheroutingindicesare
computedandappliedusingthehistoricalnetworkparameters𝜃 𝑡−Δ𝑡. Inpractice,tocircumvent
the overhead of loading model parameters twice, we fetch the data for step 𝑡 in advance at
step𝑡−Δ𝑡. We"anticipatorily"computeandcachetheroutingindicestobeusedlateratstep
𝑡,whichiswhywenamethisapproachAnticipatoryRouting. Wealsoheavilyoptimizedthis
at the infrastructure level. First, given that pre-computing the routing indices only requires
asingleforwardpassoverthedata,wecarefullyorchestratedthepipelineexecutionandthe
overlappingofcomputationwithExpertParallelism(EP)communication,successfullybounding
theadditionalwall-clocktimeoverheadofAnticipatoryRoutingtoapproximately20%. Second,
weintroducedanautomaticdetectionmechanismthattriggersashortrollbackandactivates
AnticipatoryRoutingexclusivelywhenalossspikeoccurs;afteroperatinginthismodefora
certain period, the system reverts to standard training. Ultimately, this dynamic application
allowsustoavertlossspikeswithnegligibleoveralladditionaltrainingoverhead,allwithout
compromisingmodelperformance.
SwiGLUClamping. Inpreviousliterature(Belloetal.,2017;Riviereetal.,2024), clamping
hasbeenexplicitlyutilizedtoconstrainnumericalranges,therebyenhancingtrainingstability.
Inouractualtrainingruns,weempiricallyfoundthatapplyingSwiGLUclamping(OpenAI,
2025) effectively eliminates outliers and substantially aids in stabilizing the training process,
withoutcompromisingperformance. ThroughoutthetrainingofbothDeepSeek-V4-Flashand
DeepSeek-V4-Pro,weclampedthelinearcomponentofSwiGLUtotherangeof [−10,10],while
cappingtheupperboundofthegatecomponentat10.
4.3. Evaluations
4.3.1. EvaluationBenchmarks
Fortheevaluationofthebasemodels,weconsiderbenchmarksspanningfourkeydimensions:
worldknowledge,languageunderstandingandreasoning,codingandmathematics,andlong-
contextprocessing.
WorldknowledgebenchmarksincludeAGIEval(Zhongetal.,2023),C-Eval(Huangetal.,
2023), CMMLU (Li et al., 2023) MMLU (Hendrycks et al., 2020), MMLU-Redux (Gema et al.,
2024), MMLU-Pro (Wang et al., 2024b), MMMLU (OpenAI, 2024a), MultiLoKo (Hupkes and
Bogoychev,2025),Simple-QAverified(Haasetal.,2025),SuperGPQA(Duetal.,2025),FACTS
Parametric(Chengetal.,2025),andTriviaQA(Joshietal.,2017).
LanguageunderstandingandreasoningbenchmarksincludeBigBenchHard(BBH)(Suzgun
etal.,2022),DROP(Duaetal.,2019),HellaSwag(Zellersetal.,2019),CLUEWSC(Xuetal.,2020),
andWinoGrande(Sakaguchietal.,2019).
Coding and mathematical benchmarks include BigCodeBench (Zhuo et al., 2025), Hu-
manEval(Chenetal.,2021),GSM8K(Cobbeetal.,2021),MATH(Hendrycksetal.,2021),MGSM
(Shietal.,2023),andCMath(Weietal.,2023).
LongcontextbenchmarksincludeLongBench-V2(Baietal.,2025b).
27

---

Table1 | ComparisonamongDeepSeek-V3.2-Base,DeepSeek-V4-Flash-Base,andDeepSeek-V4-
Pro-Base. Allmodelsareevaluatedinourinternalframeworkandsharethesameevaluation
setting. Scoreswithagapnotexceeding0.3areconsideredtobeatthesamelevel. Thehighest
scoreineachrowisinboldfont,andthesecondisunderlined.
DeepSeek-V3.2 DeepSeek-V4-Flash DeepSeek-V4-Pro
Benchmark(Metric) #Shots
Base Base Base
Architecture - MoE MoE MoE
#ActivatedParams - 37B 13B 49B
#TotalParams - 671B 284B 1.6T
AGIEval(EM) 0-shot 80.1 82.6 83.1
MMLU(EM) 5-shot 87.8 88.7 90.1
MMLU-Redux(EM) 5-shot 87.5 89.4 90.8
MMLU-Pro(EM) 5-shot 65.5 68.3 73.5
MMMLU(EM) 5-shot 87.9 88.8 90.3
C-Eval(EM) 5-shot 90.4 92.1 93.1
WorldKnowl.
CMMLU(EM) 5-shot 88.9 90.4 90.8
MultiLoKo(EM) 5-shot 38.7 42.2 51.1
Simple-QAverified(EM) 25-shot 28.3 30.1 55.2
SuperGPQA(EM) 5-shot 45.0 46.5 53.9
FACTSParametric(EM) 25-shot 27.1 33.9 62.6
TriviaQA(EM) 5-shot 83.3 82.8 85.6
BBH(EM) 3-shot 87.6 86.9 87.5
DROP(F1) 1-shot 88.2 88.6 88.7
Lang.&Reas. HellaSwag(EM) 0-shot 86.4 85.7 88.0
WinoGrande(EM) 0-shot 78.9 79.5 81.5
CLUEWSC(EM) 5-shot 83.5 82.2 85.2
BigCodeBench(Pass@1) 3-shot 63.9 56.8 59.2
HumanEval(Pass@1) 0-shot 62.8 69.5 76.8
GSM8K(EM) 8-shot 91.1 90.8 92.6
Code&Math
MATH(EM) 4-shot 60.5 57.4 64.5
MGSM(EM) 8-shot 81.3 85.7 84.4
CMath(EM) 3-shot 92.6 93.6 90.9
LongContext LongBench-V2(EM) 1-shot 40.2 44.7 51.5
4.3.2. EvaluationResults
InTable1,weprovideadetailedcomparisonofthebasemodelsforDeepSeek-V3.2,DeepSeek-
V4-Flash,andDeepSeek-V4-Pro,allevaluatedunderaunifiedinternalframeworkwithstrictly
consistentsettings.
Comparing DeepSeek-V4-Flash-Base with DeepSeek-V3.2-Base reveals a compelling ef-
ficiency story. Despite utilizing a substantially smaller number of both activated and total
parameters,DeepSeek-V4-Flash-BaseoutperformsDeepSeek-V3.2-Baseacrossawidearrayof
benchmarks. Thisadvantageisespeciallyevidentinworldknowledgetasksandchallenging
long-contextscenarios. Theseresultsunderscorethatarchitecturalimprovements,refineddata
quality,andtrainingoptimizationsinDeepSeek-V4-Flash-Baseyieldsuperiorperformanceeven
withamorecompactparameterbudget,effectivelysurpassingthelargerDeepSeek-V3.2-Base
onthemajorityofevaluations.
Furthermore, DeepSeek-V4-Pro-Base demonstrates a further, decisive leap in capability,
establishingnear-universaldominanceoverbothDeepSeek-V3.2-BaseandDeepSeek-V4-Flash-
Base. With improvements across almost all categories, DeepSeek-V4-Pro-Base reaches new
28

---

performance highs among DeepSeek base models on the most demanding benchmarks. On
knowledge-intensiveevaluations,itdeliversdramaticgains,whilealsosubstantiallyadvancing
long-contextunderstanding. Onmostreasoningandcodebenchmarks,DeepSeek-V4-Pro-Base
alsoexceedsbothpreviousmodels. ThiscomprehensiveupliftconfirmsDeepSeek-V4-Pro-Base
asthestrongestfoundationmodelintheDeepSeekseries,outperformingitspredecessorsacross
thespectrumofknowledge,reasoning,coding,andlong-contextcapabilities.
5. Post-Training
5.1. Post-TrainingPipeline
Followingpre-training,weconductedapost-trainingphasetoyieldthefinalmodelsofDeepSeek-
V4 series. Although the training pipeline largely mirrored that of DeepSeek-V3.2, a critical
methodological substitution was made: the mixed Reinforcement Learning (RL) stage was
entirelyreplacedbyOn-PolicyDistillation(OPD).
5.1.1. SpecialistTraining
ThedevelopmentofdomainspecialistswasconductedbyadaptingtheDeepSeek-V3.2training
pipeline. Specifically, each model was sequentially optimized through an initial fine-tuning
phaseandsubsequentReinforcementLearning(RL)guidedbydomain-specificpromptsandre-
wardsignals. FortheRLstage,weimplementedtheGroupRelativePolicyOptimization(GRPO)
algorithm,maintaininghyper-parameterscloselyalignedwithourpriorresearch(DeepSeek-AI,
2025;DeepSeek-AI,2025).
Reasoning Efforts. It is widely recognized that a model’s performance on reasoning tasks
isfundamentallygovernedbythecomputationaleffortexpended. Consequently,wetrained
distinctspecialistmodelsunderdivergentRLconfigurationstofacilitatethedevelopmentof
modelsoptimizedforvaryingreasoningcapacities. AsdetailedinTable2,DeepSeek-V4-Proand
DeepSeek-V4-Flashbothsupportthreespecificreasoningeffortmodes. Foreachmode,weapply
distinct length penalties and context windows during RL training, which results in varying
output token lengths for reasoning. To integrate these distinct reasoning modes, we utilize
specializedresponseformatsdemarcatedbythe<think>and</think>tokens. Furthermore,
for the "Think Max" mode, we prepend a specific instruction to the beginning of the system
prompttoguidethemodel’sreasoningprocess,asshowninTable3.
GenerativeRewardModel. Typically,easy-to-verifytaskscanbeeffectivelyoptimizedusing
simplerule-basedverifiersortestcases. Incontrast,hard-to-verifytaskstraditionallyrelyon
ReinforcementLearningfromHumanFeedback(RLHF),whichnecessitatesextensivehuman
annotation to train a scalar reward model. In the post-training phase of DeepSeek-V4 series,
however,wedispensewiththeseconventionalscalar-basedrewardmodels. Instead,toaddress
hard-to-verifytasks,wecuraterubric-guidedRLdataandemployaGenerativeRewardModel
(GRM)toevaluatepolicytrajectories. Crucially,weapplyRLoptimizationdirectlytotheGRM
itself. In this paradigm, the actor network natively functions as the GRM, enabling the joint
optimizationofthemodel’sevaluative(judging)proficiencyalongsideitsstandardgenerative
capabilities. Byunifyingtheseroles,themodel’sinternalreasoningcapabilitiesareinherently
fusedintoitsevaluativeprocess,resultinginhighlyrobustscoring. Furthermore,thisapproach
achievessuperiorperformancewithonlyaminimalsetofdiversehumanannotations,asthe
29

---

Table2 | Comparisonofthreereasoningmodes
Reasoning Characteristics TypicalUseCases ResponseFormat
Mode
Non-think Fast, intuitive re- Routine daily tasks, </think>summary
sponses based on emergencyreactions,
habits or simple low-riskdecisions.
rules.
ThinkHigh Conscious logical Complex problem- <think> thinking
analysis, slower but solving, planning, tokens </think>
moreaccurate. medium-risk deci- summary
sions.
ThinkMax Pushreasoningtoits Exploringthebound- 1. A special system
fullest extent. Slow aryofmodelreason- promptatthebegin-
butpowerful. ingcapability. ning.
2. <think>thinking
tokens </think>
summary
Table3 | Instructioninjectedintothesystempromptforthe"ThinkMax"mode.
InjectedInstruction
ReasoningEffort: Absolutemaximumwithnoshortcutspermitted.
You MUST be very thorough in your thinking and comprehensively decompose the
problemtoresolvetherootcause,rigorouslystress-testingyourlogicagainstallpotential
paths,edgecases,andadversarialscenarios.
Explicitlywriteoutyourentiredeliberationprocess, documentingeveryintermediate
step,consideredalternative,andrejectedhypothesistoensureabsolutelynoassumption
isleftunchecked.
modelleveragesitsownlogictogeneralizeacrosscomplextasks.
Tool-Call Schema and Special Token. Consistent with our previous version, we utilize a
dedicated <think></think> tag to delineate the reasoning path. In DeepSeek-V4 series, we
introduceanewtool-callschemathatemploysaspecial"|DSML|"tokenandutilizesanXML-
basedformatfortoolinvocations,asdemonstratedinTable4. Ourexperimentsdemonstratethat
theXMLformateffectivelymitigatesescapingfailuresandreducestool-callerrors,providinga
morerobustinterfaceformodel-toolinteractions.
InterleavedThinking. DeepSeek-V3.2introducedacontextmanagementstrategythatretains
reasoningtracesacrosstool-resultroundsbutdiscardsthemuponthearrivalofnewusermes-
sages. Whileeffective,thisstillcausedunnecessarytokenwasteincomplexagenticworkflows
— each new user turn would flush all accumulated reasoning content, forcing the model to
reconstructitsproblem-solvingstatefromscratch. Leveragingtheexpanded1M-tokencontext
30

---

Table4 | Tool-callschemaforDeepSeek-V4series.
ToolCallSchema
## Tools
You have access to a set of tools to help answer the user’s question. You can
invoke tools by writing a "<|DSML|tool_calls>" block like the following:
<|DSML|tool_calls>
<|DSML|invoke name="$TOOL_NAME">
<|DSML|parameter name="$PARAMETER_NAME" string="true|false">$PARAMETER_VALUE
</|DSML|parameter>
...
</|DSML|invoke>
<|DSML|invoke name="$TOOL_NAME2">
...
</|DSML|invoke>
</|DSML|tool_calls>
String parameters should be specified as is and set ‘string="true"‘. For all
other types (numbers, booleans, arrays, objects), pass the value in JSON
format and set ‘string="false"‘.
If thinking_mode is enabled (triggered by <think>), you MUST output your
complete reasoning inside <think>...</think> BEFORE any tool calls or
final response.
Otherwise, output directly after </think> with tool calls or final response.
### Available Tool Schemas
{Tool Definition...}
You MUST strictly follow the above definedtool name and parameter schemas to
invoke tool calls.
windowofDeepSeek-V4series,wefurtherrefinethismechanismtomaximizetheeffectiveness
ofinterleavedthinkinginagenticenvironments:
• Tool-Calling Scenarios. As illustrated in Figure 7(a), all reasoning content is fully pre-
served throughout the entire conversation. Unlike DeepSeek-V3.2, which discarded
thinkingtracesuponeachnewuserturn,DeepSeek-V4seriesretainthecompletereason-
inghistoryacrossallrounds,includingacrossusermessageboundaries. Thisallowsthe
modeltomaintainacoherent,cumulativechainofthoughtoverlong-horizonagenttasks.
• GeneralConversationalScenarios. AsillustratedinFigure7(b),theoriginalstrategyis
preserved: reasoningcontentfrompreviousturnsisdiscardedwhenanewusermessage
arrives,keepingthecontextconciseforsettingswherepersistentreasoningtracesprovide
limitedbenefit.
AswithDeepSeek-V3.2,agentframeworksthatsimulatetoolinteractionsviausermessages(e.g.,
Terminus)maynottriggerthetool-callingcontextpathandthusmaynotbenefitfromenhanced
reasoningpersistence. Wecontinuetorecommendnon-thinkmodelsforsucharchitectures.
31

---

a) Thinking with tools
b) Thinking without tools
Figure7 | ThinkingmanagementofDeepSeek-V4series.
QuickInstruction. Inchatbotscenarios,anumberofauxiliarytasks(e.g.,determiningwhether
totriggerawebsearch,intentrecognition,etc.) mustbeexecutedbeforegeneratingtheresponse.
Conventionally,thesetasksarehandledbyaseparatesmallmodel,requiringredundantprefill-
ingsinceitcannotreusetheexistingKVcache. Toovercomethislimitation,weintroduceQuick
Instruction. Weappendasetofdedicatedspecialtokensdirectlytotheinputsequence,where
eachtokencorrespondstoaspecificauxiliarytask. Bydirectlyreusingthealready-computed
KVcache,thismechanismcompletelyavoidsredundantprefillingandallowscertaintasks,such
asgeneratingsearchqueriesanddeterminingauthorityanddomain,tobeexecutedinparallel.
Consequently,thisapproachsignificantlyreducestheuser-perceivedtime-to-first-token(TTFT)
andeliminatestheengineeringoverheadofmaintaininganditeratinganextrasmallmodel. The
supportedQuickInstructiontokensaresummarizedinTable5.
5.1.2. On-PolicyDistillation
Aftertrainingmultipledomain-specificexpertsviaspecializedfine-tuningandreinforcement
learning,weemploymulti-teacherOn-PolicyDistillation(OPD)astheprimarytechniquefor
mergingexpertcapabilitiesintothefinalmodel. OPDhasemergedasaneffectivepost-training
paradigm for efficiently transferring the knowledge and capabilities of domain experts to a
single,unifiedmodel. Thisisachievedbyhavingthestudentlearnfromtheoutputdistributions
ofteachermodelsonitsowngeneratedtrajectories. Formally,givenasetof 𝑁 expertmodels
32

---

Table5 | QuickInstructionspecialtokensforauxiliarytasks.
SpecialToken Description Format
<|action|> Determineswhethertheuser ...<|User|>{prompt}<|Assistant|>
promptrequiresawebsearch <think><|action|>
orcanbeanswereddirectly.
<|title|> Generatesaconciseconversa- ...<|Assistant|>{response}
tion title after the first assis- <|end_of_sentence|><|title|>
tantresponse.
<|query|> Generatessearchqueriesfor ...<|User|>{prompt}<|query|>
theuserprompt.
<|authority|> Classifies the user prompt’s ...<|User|>{prompt}<|authority|>
demandforsourceauthorita-
tiveness.
<|domain|> Identifies the domain of the ...<|User|>{prompt}<|domain|>
userprompt.
<|extracted_url|>Determines whether each ...<|User|>{prompt}
<|read_url|> URL in the user prompt <|extracted_url|>{url}
shouldbefetchedandread. <|read_url|>
{𝜋
𝐸
1
,𝜋
𝐸
2
,...,𝜋
𝐸𝑁
},theOPDobjectivefunctionisdefinedas:
𝑁
L OPD (𝜃) = ∑︁ 𝑤 𝑖 ·D KL (cid:0)𝜋 𝜃 ∥ 𝜋 𝐸𝑖 (cid:1) . (29)
𝑖=1
Inthisformulation,𝑤 𝑖 representstheassignedweightforeachexpert,typicallydeterminedby
the relative importance of the expert. Computing the reverse KL loss D KL (cid:0)𝜋 𝜃 ∥ 𝜋 𝐸𝑖 (cid:1) requires
samplingtrainingtrajectoriesfromthestudent𝜋 𝜃 tomaintainon-policylearning. Theunderly-
inglogicensuresthattheunifiedpolicy𝜋 𝜃selectivelylearnsfromthespecializedexpertrelevant
tothecurrenttaskcontext(e.g.,aligningwiththemathematicsexpertformathreasoningtasks
andthecodingexpertforprogrammingtasks). Throughthismechanism,theknowledgefrom
physicallydistinctexpertweightsisconsolidatedintoaunifiedparameterspacevialogits-level
alignment, practically circumventing the performance degradation often encountered in tra-
ditionalweight-mergingormixedRLtechniques. Inthisstage,morethantenteachermodels
coveringvariousdomainsareemployedtodistillasinglestudentmodel.
InhandlingtheaboveOPDobjective,priorworksusuallysimplifythefull-vocabularyKL
lossintoatoken-levelKLestimateateachtokenposition,andreuseRLframeworkbyreplac-
ingsg
(cid:2)
log
𝜋𝐸𝑖 (𝑦𝑡|𝑥,𝑦<𝑡)(cid:3)
(sgrepresentsthestopgradientoperation)astheper-tokenadvantage
𝜋 𝜃(𝑦𝑡|𝑥,𝑦<𝑡)
estimate in the policy loss calculation. Although this approach is resource-efficient, it leads
to high variance in gradient estimation and often causes training instability. Therefore, we
adoptfull-vocabularylogitdistillationinourOPD.Preservingthecompletelogitdistributionin
calculatingreverseKLlossyieldsmorestablegradientestimatesandensuresfaithfuldistillation
oftheteachers’knowledge. Inthefollowingsubsection,wedescribetheengineeringeffortsthat
makefull-vocabularyOPDfeasibleatscale.
33

---

5.2. RLandOPDInfrastructures
Ourpost-traininginfrastructureisbuiltuponthescalableframeworkdevelopedforDeepSeek-
V3.2. Specifically,weintegratethesamedistributedtrainingstackdescribedinSection3.5and
the rollout engine introduced earlier for efficient auto-regressive sampling. Building on this
foundation, we introduce the following principal enhancements in the present work. These
designsenableefficientexecutionofultra-long-contextRLandOPDmergingtasksinvolving
overtendistinctteachermodels,therebysubstantiallyacceleratingtheiterationcycleformodel
releases.
5.2.1. FP4QuantizationIntegration
WeapplyFP4(MXFP4)quantizationtoacceleratebothrolloutsandallinference-onlyforward
passes,includingthoseofteacherandreferencemodels,therebyreducingmemorytrafficand
sampling latency. As detailed in Section 3.4, we directly use native FP4 weights during the
rollout and inference phases. For training steps, FP4 quantization is simulated via a lossless
FP4-to-FP8dequantizationstep,allowingseamlessreuseoftheexistingFP8mixed-precision
frameworkwithFP32masterweightsandrequiringnomodificationtothebackwardpipeline.
5.2.2. EfficientTeacherSchedulingforFull-VocabularyOPD
Our framework supports full-vocabulary On-Policy Distillation (OPD) with an effectively
unboundednumberofteachers,eachpotentiallycomprisingtrillionsofparameters. Toenable
this, all teacher weights are offloaded to a centralized distributed storage and are loaded on
demandduringtheteacherforwardpasswithZeRO-likeparametershardingtoalleviateboth
I/O and DRAM pressure. Furthermore, naively materializing logits for a vocabulary size
|𝑉| > 100k across all teachers is prohibitive, even when spooled to disk. We address this by
caching only the last-layer teacher hidden states in a centralized buffer during the forward
pass. Attrainingtime,thesecachedstatesareretrievedandpassedthroughthecorresponding
predictionheadmoduletoreconstructthefulllogitsonthefly. Thisdesignincursnegligible
recomputationoverheadwhilecompletelycircumventingthememoryburdenassociatedwith
explicitlogitsmaterialization. TomitigatetheGPUmemoryfootprintoftheteacherprediction
head,weordertrainingsamplesbyteacherindexduringdatadispatching. Thisarrangement
ensures that each distinct teacher head is loaded only once per mini-batch and that at most
oneteacherheadresidesindevicememoryatanygiventime. Allparametersandhiddenstate
loading/offloadingoperationsproceedasynchronouslyinthebackground,withoutblocking
computationonthecriticalpath. Finally,theexactKLdivergencesbetweenteacherandstudent
logitsarecomputedusingaspecializedTileLangkernel,whichacceleratesthecomputationand
curtailsdynamicmemoryallocation.
5.2.3. PreemptibleandFault-TolerantRolloutService
TomaximizeGPUresourceutilizationwhileenablingrapidhardwareprovisioningforhigh-
prioritytasks,ourGPUclusteremploysacluster-widepreemptivetaskscheduler,whereany
runningtaskmaybepreemptedatanytime. Also,hardwarefailuresareprevalentinlarge-scale
GPU clusters. To this end, we implement a preemptible and fault-tolerant LLM generation
serviceforRL/OPDrollout.
Specifically,weimplementatoken-granularWrite-AheadLog(WAL)foreachgeneration
request. Wheneveranewtokenisgeneratedforarequest,weimmediatelyappendittothat
request’sWAL.Duringpreemption,wepausetheinferenceengineandsavetheKVcacheof
34

---

unfinished requests. Upon resumption, we use the persisted WALs and saved KV cache to
continuedecoding. Evenwhenafatalhardwareerroroccurs,wecanre-runtheprefillphase
usingthepersistedtokensinWALtoreconstructtheKVcache.
Importantly,itismathematicallyincorrecttoregenerateunfinishedrequestsfromscratch,
asthisintroduceslengthbias. Becauseshorterresponsesaremorelikelytosurviveinterrup-
tion,regeneratingfromscratchmakesthemodelmorepronetoproducingshortersequences
whenever an interruption occurs. If the inference stack is batch-invariant and deterministic,
this correctness issue could also be addressed by regenerating with a consistent seed for the
pseudorandomnumbergeneratorusedinthesampler. However,thisapproachstillincursthe
extracostofre-runningthedecodingphase,makingitfarlessefficientthanourtoken-granular
WALmethod.
5.2.4. ScalingRLFrameworkforMillion-TokenContext
We introduce targeted optimizations for efficient RL and OPD on million-token sequences.
Duringtherolloutphase,weadoptapreemptibleandfault-tolerantrolloutservice,detailedin
Section5.2.3. Fortheinferenceandtrainingphase,wedecomposetherolloutdataformatinto
lightweightmetadataandheavyper-tokenfields. Duringdatadispatching,themetadataforthe
entirerolloutdatacanbeloadedtoperformglobalshufflingandpackinglayoutcomputation.
Heavy per-token fields are loaded via a shared-memory data loader to eliminate intra-node
dataredundancyandarereleasedimmediatelyuponconsumptionatthemini-batchgranularity,
substantiallyreducingbothCPUandGPUmemorypressure. Thenumberofon-devicemini-
batchesisdynamicallydeterminedbasedonworkload,allowinganefficienttrade-offbetween
computationalthroughputandI/Ooverlap.
5.2.5. SandboxInfrastructureforAgenticAI
To meet the diverse execution demands of agentic AI during post-training and evaluation,
we build a production-grade sandbox platform, DeepSeek Elastic Compute (DSec). DSec
comprisesthreeRustcomponents—theAPIgateway(Apiserver),per-hostagent(Edge),and
theclustermonitor(Watcher)—thatareinterconnectedbyacustomRPCprotocolandscale
horizontallyatopthe3FSdistributedfilesystem(DeepSeek-AI,2025). Inproduction,asingle
DSecclustermanageshundredsofthousandsofconcurrentsandboxinstances.
The design of DSec is motivated by four observations: (1) agentic workloads are highly
heterogeneous,spanninglightweightfunctioncallstofullsoftware-engineeringpipelineswith
diverseOSandsecurityrequirements;(2)environmentimagesarenumerousandlarge,yetmust
loadquicklyandsupportiterativecustomization;(3)high-densitydeploymentdemandsefficient
CPUandmemoryutilization;(4)sandboxlifecyclesmustcoordinatewithGPUtrainingsched-
ules,includingpreemptionandcheckpoint-basedresumption. Basedontheseobservations,we
elaborateonthefourcoredesignsofDSecindividuallyinthefollowing.
Four Execution Substrates Behind One Unified Interface. DSec exposes a single Python
SDK (libdsec) that abstracts four execution substrates. Function Call dispatches stateless
invocationstoapre-warmedcontainerpool,eliminatingcold-startoverhead. Containerisfully
Docker-compatible and leverages EROFS (Gao et al., 2019) on-demand loading for efficient
imageassembly. microVM,builtonFirecracker(Agacheetal.,2020),addsVM-levelisolationfor
security-sensitive,high-densitydeployments. fullVM,builtonQEMU(Bellard,2005),supports
arbitraryguestoperatingsystems. AllfourshareacommonAPIsurface—commandexecution,
35

---

filetransfer,andTTYaccess—andswitchingbetweenthemrequiresonlyaparameterchange.
FastImageLoadingviaLayeredStorage. DSecreconcilesfaststartupwithalargeandgrowing
corpusofenvironmentimagesthroughlayered,on-demandloading. Forcontainers,baseimages
andfilesystemcommitsarestoredas3FS-backedreadonlyEROFSlayersmounteddirectlyinto
overlaylowerdirs. Wekeepfilemetadatareadilyavailableonthelocaldiskatmounttime;
meanwhile, data blocks are fetched from 3FS upon request. For microVMs, DSec uses the
overlaybd (Li et al., 2020) disk format: the read-only base layer resides on 3FS for cross-
instancesharing,whilewritesgotoalocalcopy-on-writelayer. Suchsnapshotsarechainable,
facilitatingefficientversioningandmillisecond-scaleresumption.
DensityOptimizationsUnderMassiveConcurrency. Toaccommodatehundredsofthousands
of sandboxes per cluster, DSec tackles two resource bottlenecks. First, it mitigates duplicate
page-cachefootprintsinvirtualizedenvironmentsandappliesmemoryreclamationtoenable
safeovercommitment. Second,italleviatesspinlockcontentioninthecontainerruntimeand
therefore,reducesper-sandboxCPUoverhead,significantlyincreasingper-hostpackingdensity.
TrajectoryLoggingandPreemption-SafeResumption. DSecmaintainsagloballyordered
trajectorylogforeachsandbox,persistentlyrecordingeverycommandinvocationanditsresults.
The trajectory serves three purposes: (1) client fast-forwarding — when a training task is
preempted,sandboxresourcesareretainednonetheless;uponresumption,DSecreplayscached
resultsforpreviouslycompletedcommands,acceleratingtaskrecoverywhilstalsopreventing
errors from re-execution of non-idempotent operations; (2) fine-grained provenance — the
originandcorrespondingoutcomesofeachstatechangearetraceable;(3)deterministicreplay
—anyhistoricalsessioncanbefaithfullyreproducedfromitstrajectory.
5.3. StandardBenchmarkEvaluation
5.3.1. EvaluationSetup
KnowledgeandReasoning. KnowledgeandreasoningdatasetsincludeMMLU-Pro(Wang
etal.,2024b),GPQA(Reinetal.,2023),HumanLastExam(Phanetal.,2025),Simple-QAVeri-
fied(Haasetal.,2025),Chinese-SimpleQA(Heetal.,2024),LiveCodeBench-v6(Jainetal.,2024),
CodeForces(InternalBenchmark),HMMT2026Feb,Apex(Balunovic´ etal.,2025),ApexShort-
list(Balunovic´ etal.,2025),IMOAnswerBench(Luongetal.,2025),andPutnamBench(Tsoukalas
etal.,2024).
Forcode,weevaluateDeepSeek-V4seriesonLiveCodeBench-v6andaninternalCodeforces
benchmark. For Codeforces, we collect 14 Codeforces Division 1 contests comprising 114
problems(May2025-November2025). TheEloratingiscomputedasfollows. Foreachcontest,
wegenerate32candidatesolutionsperproblem. Foreachproblemindependently,wesample
10 of these solutions without replacement and arrange them in a random order to form the
submission sequence. Each submission is judged against a test suite constructed by domain
experts. ThescoreforasolvedproblemfollowsthepenaltyschemeofOpenAI(2025): themodel
receivesthemedianscoreofhumanparticipantswhosolvedthesameproblemwiththesame
numberofpriorfailedattempts. Thisyieldsatotalcontestscoreforeachsampledsubmission
sequence,whichisthenconvertedintoacontestrankandsubsequentlyintoanestimatedrating
viathestandardCodeforcesratingsystem. Thecontest-levelexpectedratingisdefinedasthe
36

---

expectationofthis estimated ratingoverallpossiblerandom selectionsandorderingsof the
10 submissions per problem. The model’s overall rating is the average of these contest-level
expectedratingsacrossall14contests.
Forreasoningandknowledgetasks,wesetthetemperatureto1.0andthecontextwindowto
8K,128K,and384KtokensfortheNon-think,High,andMaxmodes,respectively. Formathtasks
(e.g.,HMMT,IMOAnswerBench,Apex,andHLE),weevaluateusingthefollowingtemplate:
"{question}\nPlease reason step by step, and put your final answer within
\boxed{}."ForDeepSeek-V4-Pro-Maxonmathtasks,weusethefollowingtemplatetoelicit
deeper reasoning: "Solve the following problem. The problem may ask you to
prove a statement, or ask for an answer. If finding an answer is required,
you should come up with the answer, and your final solution should also be
a rigorous proof of that answer being valid.\n\n{question}".
For formal math tasks, we evaluate in an agentic setting on Lean v4.28.0-rc1 (Moura and
Ullrich,2021),withaccesstotheLeancompilerandasemantictacticsearchengine,running
up to 500 tool calls with max reasoning effort. In addition, we evaluate a more compute-
intensivepipelineinwhichcandidatenatural-languagesolutionsarefirstgeneratedandfiltered
byself-verification(Shaoetal.,2025),andtheretainedsolutionsarethenprovidedasguidance
to a formal agent for proving the corresponding Lean statement. This design uses informal
reasoningtoimproveexplorationwhilepreservingstrictcorrectnessthroughformalverification.
A submission is counted as correct only if the strict verifier Comparator accepts it for both
settings.
WehaveleftsomeentriesblankforK2.6andGLM-5.1,astheirAPIsweretoobusytoreturn
responsestoourqueries.
1M-TokenContext. SinceDeepSeek-V4seriessupports1M-tokencontexts,weevaluatemodel
performance in a long context scenario by selecting OpenAI MRCR (OpenAI, 2024b) and
CorpusQA(Luetal.,2026)asthebenchmarks. Were-evaluateClaudeOpus4.6andGemini3.1
Proonthesetaskswiththegoalofstandardizingtheconfigurationacrossallmodels. Wedid
notevaluateGPT-5.4becauseitsAPIfailedtorespondtoalargeportionofourqueries.
Agent. AgentdatasetsincludeTerminalBench2.0(Merrilletal.,2026),SWE-Verified(OpenAI,
2024e),SWEMultilingual(Yangetal.,2025),SWE-Pro(Dengetal.,2025),BrowseComp(Wei
etal.,2025),thepublicevaluationsetofMCPAtlas(Bandietal.,2026),GDPval-AA(AA,2025;
Patwardhanetal.,2025),andTool-Decathlon(Lietal.,2025).
Forcodeagenttasks(SWE-Verified,Terminal-Bench,SWE-Pro,SWEMultilingual),weeval-
uateDeepSeek-V4seriesusinganinternallydevelopedevaluationframework. Thisframework
provides a minimal set of tools — a bash tool and a file-edit tool. The maximum number of
interactionstepsissetto500,andthemaximumcontextlengthissetto512Ktokens. Regarding
Terminal-Bench2.0,weacknowledgetheenvironment-relatedissuesnotedbyGLM-5.1. Never-
theless,wereportourperformanceontheoriginalTerminal-Bench2.0datasetforconsistency.
OntheTerminal-Bench2.0Verifiedsubset,DeepSeek-V4-Proachievesascoreofapproximately
72.0.
Forsearchagenttasks(BrowseComp,HLEw/tool),wealsouseanin-househarnesswith
websearchandPythontool,andsetmaximuminteractionstepsto500andthemaximumcontext
length to 512K tokens. For BrowseComp, we use the same discard-all context management
strategyasDeepSeek-V3.2(DeepSeek-AI,2025).
37

---

5.3.2. EvaluationResults
Table6 | ComparisonbetweenDeepSeek-V4-Pro-Maxandclosed/opensourcemodels. "Max",
"xHigh", and "High" denote reasoning effort. The best results are highlighted in bold; the
second-bestresultsareunderlined.
Opus-4.6GPT-5.4Gemini-3.1-Pro K2.6 GLM-5.1 DS-V4-Pro
Benchmark(Metric)
Max xHigh High ThinkingThinking Max
gninosaeR&egdelwonK
MMLU-Pro(EM) 89.1 87.5 91.0 87.1 86.0 87.5
SimpleQA-Verified(Pass@1) 46.2 45.3 75.6 36.9 38.1 57.9
Chinese-SimpleQA(Pass@1) 76.4 76.8 85.9 75.9 75.0 84.4
GPQADiamond(Pass@1) 91.3 93.0 94.3 90.5 86.2 90.1
HLE(Pass@1) 40.0 39.8 44.4 36.4 34.7 37.7
LiveCodeBench(Pass@1) 88.8 - 91.7 89.6 - 93.5
Codeforces(Rating) - 3168 3052 - - 3206
HMMT2026Feb(Pass@1) 96.2 97.7 94.7 92.7 89.4 95.2
IMOAnswerBench(Pass@1) 75.3 91.4 81.0 86.0 83.8 89.8
Apex(Pass@1) 34.5 54.1 60.9 24.0 11.5 38.3
ApexShortlist(Pass@1) 85.9 78.1 89.1 75.5 72.4 90.2
gnoL MRCR1M(MMR) 92.9 - 76.3 - - 83.5
CorpusQA1M(ACC) 71.7 - 53.8 - - 62.0
citnegA
TerminalBench2.0(Acc) 65.4 75.1 68.5 66.7 63.5 67.9
SWEVerified(Resolved) 80.8 - 80.6 80.2 - 80.6
SWEPro(Resolved) 57.3 57.7 54.2 58.6 58.4 55.4
SWEMultilingual(Resolved) 77.5 - - 76.7 73.3 76.2
BrowseComp(Pass@1) 83.7 82.7 85.9 83.2 79.3 83.4
HLEw/tools(Pass@1) 53.1 52.0 51.6 54.0 50.4 48.2
GDPval-AA(Elo) 1619 1674 1314 1482 1535 1554
MCPAtlasPublic(Pass@1) 73.8 67.2 69.2 66.6 71.8 73.6
Toolathlon(Pass@1) 47.2 54.6 48.8 50.0 40.7 51.8
ThecomparisonofDeepSeek-V4-Pro-Maxandotherclosed/opensourcemodelsispresented
inTable6. Also,weevaluatedifferentmodesofDeepSeek-V4-FlashandDeepSeek-V4-Proand
showtheresultsinTable7.
Knowledge. Intheevaluationofgeneralworldknowledge,DeepSeek-V4-Pro-Max,themax-
imum reasoning effort mode of DeepSeek-V4-Pro, establishes a new state-of-the-art among
open-sourcelargelanguagemodels. AsdemonstratedbytheSimpleQA-Verified,DeepSeek-V4-
Pro-Maxsignificantlyoutperformsallexistingopen-sourcebaselinesbyamarginof20absolute
percentage points. Despite these advances, it currently trails the leading proprietary model,
Gemini-3.1-Pro. Inthedomainofeducationalknowledgeandreasoning,DeepSeek-V4-Pro-Max
marginallyoutperformsKimiandGLMacrosstheMMLU-Pro,GPQA,andHLEbenchmarks,
althoughitlagsbehindleadingproprietarymodels. Broadly,DeepSeek-V4-Pro-Maxmarksa
significantmilestoneinenhancingtheworldknowledgecapabilitiesofopen-sourcemodels.
Inaddition,asignificantperformancegapexistsbetweenDeepSeek-V4-FlashandDeepSeek-
V4-Pro on knowledge-based tasks; this is anticipated, as larger parameter counts facilitate
greaterknowledgeretentionduringpre-training. Notably,bothmodelsdemonstrateimproved
resultsonknowledgebenchmarkswhenallocatedhigherreasoningeffort.
38

---

Table7 | ComparisonamongdifferentsizesandmodesofDeepSeek-V4series. "Non-Think",
"High",and"Max"denotereasoningeffort.
DeepSeek-V4-Flash DeepSeek-V4-Pro
Benchmark(Metric)
Non-Think High Max Non-Think High Max
gninosaeR&egdelwonK
MMLU-Pro(EM) 83.0 86.4 86.2 82.9 87.1 87.5
SimpleQA-Verified(Pass@1) 23.1 28.9 34.1 45.0 46.2 57.9
Chinese-SimpleQA(Pass@1) 71.5 73.2 78.9 75.8 77.7 84.4
GPQADiamond(Pass@1) 71.2 87.4 88.1 72.9 89.1 90.1
HLE(Pass@1) 8.1 29.4 34.8 7.7 34.5 37.7
LiveCodeBench(Pass@1-COT) 55.2 88.4 91.6 56.8 89.8 93.5
Codeforces(Rating) - 2816 3052 - 2919 3206
HMMT2026Feb(Pass@1) 40.8 91.9 94.8 31.7 94.0 95.2
IMOAnswerBench(Pass@1) 41.9 85.1 88.4 35.3 88.0 89.8
Apex(Pass@1) 1.0 19.1 33.0 0.4 27.4 38.3
ApexShortlist(Pass@1) 9.3 72.1 85.7 9.2 85.5 90.2
gnoL MRCR1M(MMR) 37.5 76.9 78.7 44.7 83.3 83.5
CorpusQA1M(ACC) 15.5 59.3 60.5 35.6 56.5 62.0
citnegA
TerminalBench2.0(Acc) 49.1 56.6 56.9 59.1 63.3 67.9
SWEVerified(Resolved) 73.7 78.6 79.0 73.6 79.4 80.6
SWEPro(Resolved) 49.1 52.3 52.6 52.1 54.4 55.4
SWEMultilingual(Resolved) 69.7 70.2 73.3 69.8 74.1 76.2
BrowseComp(Pass@1) - 53.5 73.2 - 80.4 83.4
HLEw/tools(Pass@1) - 40.3 45.1 - 44.7 48.2
MCPAtlasPublic(Pass@1) 64.0 67.4 69.0 69.4 74.2 73.6
GDPval-AA(Elo) - - 1395 - - 1554
Toolathlon(Pass@1) 40.7 43.5 47.8 46.3 49.0 51.8
Reasoning. DeepSeek-V4-Pro-Maxoutperformsallprioropenmodelsacrossreasoningbench-
marks,andmatchesstate-of-the-artclosedmodelsonmanymetrics,whilethesmallerDeepSeek-
V4-Flash-Max also surpasses the previous best open-source model, K2.6-Thinking, on code
andmathreasoningtasks. Meanwhile,DeepSeek-V4-ProandDeepSeek-V4-Flashexcelincod-
ing competitions. According to our evaluation, their performance is comparable to GPT-5.4,
making this the first time an open model has matched a closed model on this task. On the
Codeforcesleaderboard,DeepSeek-V4-Pro-Maxcurrentlyranks23rdamonghumancandidates.
DeepSeek-V4alsodemonstratesstrongperformanceonformalmathematicaltaskunderboth
agentic and compute-intensive settings. Under an agentic setup, it achieves state-of-the-art
results,showninFigure8,outperformingpriormodelssuchasSeedProver(Chenetal.,2025).
Withamorecompute-intensivepipeline,performancefurtherimproves,surpassingsystems
includingAristotle(Achimetal.,2025)andmatchingthebestknownresultsunderthissetting.
Agent. TheDeepSeek-V4seriesdemonstratesstrongagentperformanceinevaluations. For
codeagenttasks,DeepSeek-V4-ProachievesresultscomparabletoK2.6andGLM-5.1,though
all these open models still lag behind their closed-source counterparts. DeepSeek-V4-Flash
underperformsDeepSeek-V4-Prooncodingtasks,particularlyonTerminalBench2.0. Asimilar
trend is observed across other agent evaluations. It is worth noting that DeepSeek-V4-Pro
performswellonMCPAtlasandToolathlon—twoevaluationtestsetsthatincludeawiderange
oftoolsandMCPservices—indicatingthatourmodelhasexcellentgeneralizationcapability
anddoesnotperformwellonlyoninternalframeworks.
39

---

PracticalRegime FrontierRegime
Putnam-200Pass@8withminimaltools Putnam-2025withhybridformal-informal
andboundedsampling. reasoningandsubstantialcomputescaling.
Seed-1.5-Prover 26.50 Aristotle 100/120
Gemini-3-Pro 26.50 Seed-1.5-Prover 110/120
Seed-2.0-Pro 35.50 Axiom 120/120
DeepSeek-V4-Flash-Max 81.00 DeepSeek-V4 120/120
Figure 8 | Formal reasoning under practical and frontier regimes. Left: Putnam-200 Pass@8
evaluatesafixedrandomsubsetofPutnamBench(Tsoukalasetal.,2024)followingthesetup
introducedbySeed-Prover;allmodelsaretestedonthesameproblemset. WefollowtheSeed-
Proverprotocolbutreplaceproprietarysearchtoolswiththeopen-sourceLeanExplore(Asher,
2025),yieldingalightweightsettingwithminimalagenttoolsandboundedsampling. Right:
Putnam-2025probesthefrontierofmathematicalreasoninginascaledhybridformal-informal
regime, where informal reasoning is combined with formal verification to expose gaps and
improverigor;DeepSeek-V4reachesaproof-perfect120/120.
1.0
0.8
0.6
0.4
0.2
0.0
8k 16k 32k 64k 128k 256k 512k 1024k
Input Tokens
RMM
egarevA
MRCR 8-needle
0.94 0.90 0.90 0.92
0.85
0.82
0.91
0.87 0.87 0.84 0.85
0.66
0.76
0.59
0.60
0.49
DeepSeek-V4-Pro-Max
DeepSeek-V4-Flash-Max
Figure9 | DeepSeek-V4seriesperformanceontheMRCRtask.
1M-TokenContext. DeepSeek-V4-ProoutperformsGemini-3.1-ProontheMRCRtask,which
measuresin-contextretrieval,butremainsbehindClaudeOpus4.6. AsillustratedinFigure9,
retrievalperformanceremainshighlystablewithina128Kcontextwindow. Whileaperformance
degradation becomes visible beyond the 128K mark, the model’s retrieval capabilities at 1M
tokensremainremarkablystrongcomparedtobothproprietaryandopen-sourcecounterparts.
UnlikeMRCR,CorpusQAissimilartorealscenarios. Theevaluationresultsalsoindicatethat
DeepSeek-V4-ProisbetterthanGemini-3.1-Pro.
ReasoningEffort. AsshowninTable7,theMaxmode,whichemployslongercontextsand
reduced length penalties in RL, outperforms the High mode on the most challenging tasks.
Figure10presentsacomparisonofperformanceandcostamongDeepSeek-V4-Pro,DeepSeek-
V4-Flash,andDeepSeek-V3.2onrepresentativereasoningandagentictasks. Byscalingtest-time
compute,DeepSeek-V4seriesachievesubstantialimprovementsoverthepredecessor. Further-
more,onreasoningtaskslikeHLE,DeepSeek-V4-Prodemonstrateshighertokenefficiencythan
40

---

40
35
30
25
20
15
10
5
0 20k 40k 60k 80k
Total Tokens
)%(
1@ssaP
HLE
Max 70
High
Max
Speciale 60
Think High
50
40
DeepSeek-V4-Pro
DeepSeek-V4-Flash
NNoonnee DeepSeek-V3.2
None 30
20k 30k 40k 50k
Total Tokens
)%(
1@ssaP
TerminalBench 2.0
Max
High
None
High Max
Think
None
DeepSeek-V4-Pro
DeepSeek-V4-Flash
None DeepSeek-V3.2
Figure 10 | HLE and Terminal Bench 2.0 performance by reasoning effort. “None” indicates
Non-thinkmode,and“Speciale”indicatesDeepSeek-V3.2-Specialemodel.
DeepSeek-V3.2.
5.4. PerformanceonReal-WorldTasks
Standardized benchmarks often struggle to capture the complexities of diverse, real-world
tasks,creatingagapbetweentestresultsandactualuserexperience. Tobridgethis,wehave
developedproprietaryinternalmetricsthatprioritizereal-worldusagepatternsovertraditional
benchmarks. This approach ensures that our optimizations translate into tangible benefits.
OurevaluationframeworkspecificallytargetstheprimaryusecasesoftheDeepSeekAPIand
Chatbot,aligningmodelperformancewithpracticaldemands.
5.4.1. ChineseWriting
One of the primary use cases for DeepSeek is Chinese writing. We conducted a rigorous
evaluationonfunctionalwritingandcreativewriting. Table12presentsapairwisecomparison
betweenDeepSeek-V4-ProandGemini-3.1-Proonfunctionalwritingtasks. Thesetasksconsist
of common daily writing queries, where prompts are typically concise and straightforward.
Gemini-3.1-Prowasselectedasthebaseline,asitstandsasthetop-performingexternalmodel
forChinesewritinginourevaluations. TheresultsindicatethatDeepSeek-V4-Prooutperforms
thebaselinewithanoverallwinrateof62.7%versus34.1%;thisisprimarilybecauseGemini
occasionallyallowsitsinherentstylisticpreferencestooverridetheuser’sexplicitrequirements
inChinesewritingscenarios.
Table 13 presents the creative writing comparison, which is evaluated along two axes:
instructionfollowingandwritingquality. ComparedwithGemini-3.1-Pro,DeepSeek-V4-Pro
achievesa60.0%winrateininstructionfollowingand77.5%inwritingquality,demonstrating
a marginal improvement in instruction following and a substantial gain in writing quality.
AlthoughDeepSeek-V4-Proyieldssuperiorresultsinaggregateusercaseanalysis,anevaluation
restricted to the most challenging prompts — specifically those involving high-complexity
constraints or multi-turn scenarios — reveals that Claude Opus 4.5 retains a performance
advantageoverDeepSeek-V4-Pro. AsshowninTable14,ClaudeOpus4.5achievesa52.0%win
rateversus45.9%.
41

---

5.4.2. Search
Search-augmented question answering is a core capability of the DeepSeek chatbot. On the
DeepSeekwebandapp,the"non-think"modeemploysRetrieval-AugmentedSearch(RAG),
whereasthe"thinking"modeutilizesagenticsearch.
RetrievalAugmentedSearch. WeconductedapairwiseevaluationcomparingDeepSeek-V4-
ProandDeepSeek-V3.2acrossbothobjectiveandsubjectiveQ&Acategories. Aspresentedin
Table11,DeepSeek-V4-ProoutperformsDeepSeek-V3.2byasubstantialmargin,demonstrating
a consistent advantage across both categories. The most pronounced gains are observed in
single-value search and planning & strategy tasks, suggesting that DeepSeek-V4-Pro excels
atlocating precise factualanswers and synthesizingstructuredplans from retrievedcontext.
However,DeepSeek-V3.2remainsrelativelycompetitiveoncomparisonandrecommendation
tasks,indicatingpotentialroomforimprovementforDeepSeek-V4-Proinscenariosrequiring
balanced,multi-perspectivereasoningoversearchresults.
Agentic Search. Unlike standard RAG, agentic search empowers the model to iteratively
invokesearchandfetchtoolsperquery,significantlyenhancingoverallsearchperformance. For
thethinkingmodeinDeepSeek-Chat,weoptimizedtheagenticsearchfunctiontomaximize
responseaccuracywithinapredefined"thinkingbudget". AsshowninTable9,agenticsearch
consistentlyoutperformsRAG,particularlyoncomplextasks. Furthermore, itscostremains
highlyefficient,withagenticsearchbeingonlymarginallymoreexpensivethanstandardRAG
(seeTable10).
5.4.3. White-CollarTask
Torigorouslyevaluatethemodel’sutilityinsophisticatedenterpriseproductivityscenarios,we
constructedacomprehensivesuiteof30advancedChineseprofessionaltasks. Theseworkflows
deliberatelyencompasshigh-levelcognitivedemands,includingin-depthinformationanalysis,
comprehensive document generation, and nuanced document editing, spanning a diverse
spectrumof13criticalindustries(e.g.,finance,education,law,andtechnology). Theevaluation
wasconductedwithinanin-houseagentharnessequippedwithbasictools,includingBashand
websearch.
Giventheopen-endednatureofthesetasks,automatedmetricsusuallyfallshortincapturing
thenuancesofahigh-qualityresponse. Therefore,weconductedhumanevaluationstocompare
theperformanceofDeepSeek-V4-Pro-MaxagainstOpus-4.6-Max. Annotatorsblindlyassessed
themodeloutputsacrossfourdimensions:
• TaskCompletion: Whetherthecoreproblemwassuccessfullyresolved.
• InstructionFollowing: Adherencetospecificconstraintsanddirectives.
• ContentQuality: Factualaccuracy,logicalcoherence,andprofessionaltone.
• FormattingAesthetics: Layoutreadabilityandvisualpresentation.
AsillustratedinFigure11,DeepSeek-V4-Pro-MaxoutperformsOpus-4.6-Maxondiverse
Chinesewhite-collartasks,achievinganimpressivenon-lossrateof63%,anddemonstrating
consistentadvantagesacrossanalysis,generation,andeditingtasks. Thedetaileddimension
scores shown in Figure 12 highlight the model’s primary strengths in Task Completion and
42

---

Content Quality. Specifically, DeepSeek-V4-Pro-Max proactively anticipates implicit user in-
tentsbyfrequentlyprovidingsupplementaryinsightsandself-verificationsteps. Italsoexcels
in long-form generation, delivering in-depth, coherent narratives rather than relying on the
overlysimplisticbulletpointsfrequentlyproducedbyOpus-4.6-Max. Additionally,themodel
strictlyconformstoformalprofessionalconventions,suchasstandardizedChinesehierarchical
numbering. However, in terms of Instruction Following, it occasionally overlooks specific
formatting constraints and slightly trails Opus. Furthermore, the model is less proficient at
condensing extensive text inputs into succinct summaries. Finally, its Formatting Aesthetics
stillhavesubstantialroomforimprovementregardingtheoverallvisualdesignofpresentation
slides. Figure 13, 14, and 15 present several test cases; due to the extensive length of certain
outputs,onlypartialpagesaredisplayed.
Win Rate: DeepSeek-V4-Pro-Max vs Opus-4.6-Max
analysis 55.0% 8.0% 37.0% 100
95
generation 52.0% 10.0% 38.0% 90
85
editing 47.0% 18.0% 35.0%
80
overall 53.0% 10.0% 37.0% 75
70
0% 20% 40% Proportion 60% 80% 100% Task Completio
In
n struction Following Content Qua
F
li
o
ty rmatting Aesthetics Overall
Win Tie Lose
Figure11 | Win-ratecomparisonacrossanaly-
sis,generation,editingtasks,andtheoverall
performance.
erocS
Score: DeepSeek-V4-Pro-Max vs Opus-4.6-Max
98.32
96.68
88.88
87.76 86.52
83.32 84.06
78.00
76.68
72.68
DeepSeek-V4-Pro-Max Opus-4.6-Max
Figure12 | Detaileddimensionscoresinclud-
ingTaskCompletion,ContentQuality,Format-
tingAesthetics,andInstructionFollowing.
Figure13 | Exampleoutputofataskwhichrequiresdraftingajointmarketingproposalfora
popularbubbleteabrandandtheBeijingSubway.
43

---

5.4.4. CodeAgent
Tobenchmarkourcodingagentcapability,wecuratetasksfromrealinternalR&Dworkloads
Wecollect∼200challengingtasksfrom50+internalengineers,spanningfeaturedevelopment,
bug fixing, refactoring, and diagnostics across diverse technology stacks including PyTorch,
CUDA,Rust,andC++. Eachtaskisaccompaniedbyitsoriginalrepository,thecorresponding
executionenvironment,andhuman-annotatedscoringrubrics;afterrigorousqualityfiltering,
30tasksareretainedastheevaluationset. AsshowninTable8,DeepSeek-V4-Prosignificantly
outperformsClaudeSonnet4.5andapproachesthelevelofClaudeOpus4.5.
Table8 | ComparisononR&DCodingBenchmark(externalmodelsincludedstrictlyforevalua-
tionpurposes).
Opus4.5 Opus4.6
Model Haiku4.5 Sonnet4.5 DeepSeek-V4-Pro-Max Opus4.5 Thinking Thinking
PassRate(%) 13 47 67 70 73 80
InasurveyaskingDeepSeekdevelopersandresearchers(𝑁 =85)—allwithexperienceof
usingDeepSeek-V4-Proforagenticcodingintheirdailywork—whetherDeepSeek-V4-Prois
readytoserveastheirdefaultandprimarycodingmodelcomparedtootherfrontiermodels,52%
saidyes,39%leanedtowardyes,andfewerthan9%saidno. RespondentsfindDeepSeek-V4-Pro
todeliversatisfactoryresultsacrossmosttasks,butnotetrivialmistakes,misinterpretationof
vagueprompts,andoccasionalover-thinking.
6. Conclusion, Limitations, and Future Directions
Inthiswork,wepresentapreviewversionofDeepSeek-V4series,aimingatnext-generation
largelanguagemodelsthatbreaktheefficiencybarrierofultra-long-contextprocessing. Bycom-
biningahybridattentionarchitecturethatintegratesCSAandHCA,DeepSeek-V4seriesachieve
adramaticleapinlong-sequenceefficiency. Thearchitecturalinnovations,togetherwithexten-
siveinfrastructureoptimization,enableefficientnativesupportformillion-tokencontextsand
establishanecessaryfoundationforfuturetest-timescaling,long-horizontasks,andemerging
paradigmssuchasonlinelearning. EvaluationresultsdemonstratethatDeepSeek-V4-Pro-Max,
themaximumreasoningeffortmodeofDeepSeek-V4-Pro,redefinesthestate-of-the-artforopen
models. It substantially outperforms prior open-source models on knowledge benchmarks,
achievessuperiorreasoningperformanceclosetothefrontierproprietarymodels,anddelivers
competitiveagentcapabilities. Meanwhile,DeepSeek-V4-Flash-Maxattainscomparablereason-
ingperformancetoleadingclosedmodelswhilemaintainingahighlycost-efficientarchitecture.
WebelieveDeepSeek-V4seriesusherinaneweraofmillion-lengthcontextsforopenmodels
andpavethewaytowardbetterefficiency,scale,andintelligence.
Inpursuitofextremelong-contextefficiency,DeepSeek-V4seriesadoptedaboldarchitec-
tural design. To minimize risk, we retained many preliminarily validated components and
tricks, which, while effective, made the architecture relatively complex. In future iterations,
wewillcarryoutmorecomprehensiveandprincipledinvestigationstodistillthearchitecture
down to its most essential designs, making it more elegant without sacrificing performance.
Meanwhile,althoughAnticipatoryRoutingandSwiGLUClampinghavebeenproveneffective
inmitigatingtraininginstabilities,theirunderlyingprinciplesremaininsufficientlyunderstood.
Wewillactivelystudyfoundationalproblemsontrainingstabilityandstrengtheninternalmetric
monitoring,aimingforamoreprincipledandpredictiveapproachtostablelarge-scaletraining.
44

---

Inaddition,beyondtheMoEandsparseattentionarchitecture,wewillalsoproactivelyexplore
model sparsity along new dimensions — such as more sparse embedding modules (Cheng
etal.,2026)—tofurtherimprovecomputationalandmemoryefficiencywithoutcompromising
capability. We will also continuously investigate low-latency architectures and system tech-
niquestomakelong-contextdeploymentandinteractionmoreresponsive. Furthermore,we
recognizetheimportanceandpracticalvalueoflong-horizon,multi-roundagentictasks,and
will continue to iterate and explore in this direction. We are also working on incorporating
multimodal capabilities to our models. Finally, we are committed to developing better data
curationandsynthesisstrategiestoconsistentlyenhancemodelintelligence,robustness,and
practicalusabilityacrossanincreasinglybroadrangeofscenariosandtasks.
References
AA. Gdpval-aaleaderboard,2025. URLhttps://artificialanalysis.ai/methodolog
y/intelligence-benchmarking#gdpval-aa.
T.Achim,A.Best,A.Bietti,K.Der,M.Fédérico,S.Gukov,D.Halpern-Leistner,K.Henningsgard,
Y. Kudryashov, A. Meiburg, et al. Aristotle: Imo-level automated theorem proving. arXiv
preprintarXiv:2510.01346,2025.
A.Agache,M.Brooker,A.Florescu,A.Iordache,A.Liguori,R.Neugebauer,P.Piwonka,and
D.-M.Popa. Firecracker: lightweightvirtualizationforserverlessapplications. InProceedings
ofthe17thUsenixConferenceonNetworkedSystemsDesignandImplementation,NSDI’20,
page419–434,USA,2020.USENIXAssociation. ISBN9781939133137.
O.J.Aimuyo,B.Oh,andR.Singh. Flashmoe: Fastdistributedmoeinasinglekernel. Advances
inNeuralInformationProcessingSystems,2025. URLhttps://neurips.cc/virtual/2
025/poster/119124.
J.Ainslie,J.Lee-Thorp,M.deJong,Y.Zemlyanskiy,F.Lebrón,andS.Sanghai. Gqa: Training
generalizedmulti-querytransformermodelsfrommulti-headcheckpoints. arXivpreprint
arXiv:2305.13245,2023.
J.Asher. LeanExplore: AsearchengineforLean4declarations,2025. URLhttps://arxiv.or
g/abs/2506.11085.
Y.Bai,Y.Bao,G.Chen,J.Chen,N.Chen,R.Chen,Y.Chen,Y.Chen,Y.Chen,Z.Chen,J.Cui,
H.Ding,M.Dong,A.Du,C.Du,D.Du,Y.Du,Y.Fan,Y.Feng,K.Fu,B.Gao,H.Gao,P.Gao,
T.Gao,X.Gu,L.Guan,H.Guo,J.Guo,H.Hu,X.Hao,T.He,W.He,W.He,C.Hong,Y.Hu,
Z.Hu,W.Huang,Z.Huang,Z.Huang,T.Jiang,Z.Jiang,X.Jin,Y.Kang,G.Lai,C.Li,F.Li,
H.Li,M.Li,W.Li,Y.Li,Y.Li,Z.Li,Z.Li,H.Lin,X.Lin,Z.Lin,C.Liu,C.Liu,H.Liu,J.Liu,
J.Liu,L.Liu,S.Liu,T.Y.Liu,T.Liu,W.Liu,Y.Liu,Y.Liu,Y.Liu,Y.Liu,Z.Liu,E.Lu,L.Lu,
S.Ma,X.Ma,Y.Ma,S.Mao,J.Mei,X.Men,Y.Miao,S.Pan,Y.Peng,R.Qin,B.Qu,Z.Shang,
L.Shi,S.Shi,F.Song,J.Su,Z.Su,X.Sun,F.Sung,H.Tang,J.Tao,Q.Teng,C.Wang,D.Wang,
F. Wang, and H. Wang. Kimi K2: open agentic intelligence. CoRR, abs/2507.20534, 2025a.
URLhttps://doi.org/10.48550/arXiv.2507.20534.
Y.Bai,S.Tu,J.Zhang,H.Peng,X.Wang,X.Lv,S.Cao,J.Xu,L.Hou,Y.Dong,etal. Longbench
v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics
(Volume1: LongPapers),pages3639–3664,2025b.
45

---

M.Balunovic´,J.Dekoninck,I.Petrov,N.Jovanovic´,andM.Vechev. Matharena: Evaluatingllms
on uncontaminated math competitions. Proceedings of the Neural Information Processing
SystemsTrackonDatasetsandBenchmark,2025.
C.Bandi,B.Hertzberg,G.Boo,T.Polakam,J.Da,S.Hassaan,M.Sharma,A.Park,E.Hernandez,
D. Rambado, et al. Mcp-atlas: A large-scale benchmark for tool-use competency with real
mcpservers. arXivpreprintarXiv:2602.00933,2026.
F. Bellard. Qemu, a fast and portable dynamic translator. In Proceedings of the Annual
ConferenceonUSENIXAnnualTechnicalConference,ATEC’05,page41,USA,2005.USENIX
Association.
I.Bello,H.Pham,Q.V.Le,M.Norouzi,andS.Bengio. Neuralcombinatorialoptimizationwith
reinforcementlearning,2017. URLhttps://openreview.net/forum?id=rJY3vK9eg.
J. Chen, W.Chen, J. Du, J. Hu, Z. Jiang, A. Jie, X. Jin, X.Jin, C. Li, W. Shi, Z. Wang, M.Wang,
C. Wei, S. Wei, H. Xin, F. Yang, W. Gao, Z. Yuan, T. Zhan, Z. Zheng, T. Zhou, and T. H.
Zhu. Seed-prover 1.5: Mastering undergraduate-level theorem proving via learning from
experience,2025. URLhttps://arxiv.org/abs/2512.17260.
M.Chen,J.Tworek,H.Jun,Q.Yuan,H.P.deOliveiraPinto,J.Kaplan,H.Edwards,Y.Burda,
N.Joseph,G.Brockman,A.Ray,R.Puri,G.Krueger,M.Petrov,H.Khlaaf,G.Sastry,P.Mishkin,
B.Chan,S.Gray,N.Ryder,M.Pavlov,A.Power,L.Kaiser,M.Bavarian,C.Winter,P.Tillet,
F.P.Such,D.Cummings,M.Plappert,F.Chantzis,E.Barnes,A.Herbert-Voss,W.H.Guss,
A.Nichol,A.Paino,N.Tezak,J.Tang,I.Babuschkin,S.Balaji,S.Jain,W.Saunders,C.Hesse,
A.N.Carr,J.Leike,J.Achiam,V.Misra,E.Morikawa,A.Radford,M.Knight,M.Brundage,
M.Murati,K.Mayer,P.Welinder,B.McGrew,D.Amodei,S.McCandlish,I.Sutskever,and
W.Zaremba. Evaluatinglargelanguagemodelstrainedoncode. CoRR,abs/2107.03374,2021.
URLhttps://arxiv.org/abs/2107.03374.
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze,
C. Guestrin, and A. Krishnamurthy. TVM: An automated End-to-End optimizing com-
piler for deep learning. In 13th USENIX Symposium on Operating Systems Design and
Implementation (OSDI 18), pages 578–594, Carlsbad, CA, Oct. 2018. USENIX Association.
ISBN978-1-939133-08-3. URLhttps://www.usenix.org/conference/osdi18/prese
ntation/chen.
A.Cheng,A.Jacovi,A.Globerson,B.Golan,C.Kwong,C.Alberti,C.Tao,E.Ben-David,G.S.
Tomar,L.Haas,etal. Thefactsleaderboard: Acomprehensivebenchmarkforlargelanguage
modelfactuality. arXivpreprintarXiv:2512.10791,2025.
X.Cheng,W.Zeng,D.Dai,Q.Chen,B.Wang,Z.Xie,K.Huang,X.Yu,Z.Hao,Y.Li,H.Zhang,
H.Zhang,D.Zhao,andW.Liang. Conditionalmemoryviascalablelookup: Anewaxisof
sparsityforlargelanguagemodels. CoRR,abs/2601.07372,2026. doi: 10.48550/ARXIV.2601.
07372. URLhttps://doi.org/10.48550/arXiv.2601.07372.
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek,
J.Hilton, R.Nakano, etal. Trainingverifierstosolvemathwordproblems. arXivpreprint
arXiv:2110.14168,2021.
D.Dai,C.Deng,C.Zhao,R.X.Xu,H.Gao,D.Chen,J.Li,W.Zeng,X.Yu,Y.Wu,Z.Xie,Y.K.
Li,P.Huang,F.Luo,C.Ruan,Z.Sui,andW.Liang. Deepseekmoe: Towardsultimateexpert
specialization in mixture-of-experts language models. CoRR, abs/2401.06066, 2024. URL
https://doi.org/10.48550/arXiv.2401.06066.
46

---

T.Dao,D.Haziza,F.Massa,andG.Sizov. Flash-decodingforlong-contextinference,2023. URL
https://pytorch.org/blog/flash-decoding/.
L. De Moura and N. Bjørner. Z3: an efficient smt solver. In Proceedings of the Theory
and Practice of Software, 14th International Conference on Tools and Algorithms for the
ConstructionandAnalysisofSystems,TACAS’08/ETAPS’08,page337–340,Berlin,Heidel-
berg,2008.Springer-Verlag. ISBN3540787992.
DeepSeek-AI. Deepseek-coder-v2: Breakingthebarrierofclosed-sourcemodelsincodeintelli-
gence. CoRR,abs/2406.11931,2024. URLhttps://doi.org/10.48550/arXiv.2406.11
931.
DeepSeek-AI. Deepseek-v3technicalreport. CoRR,abs/2412.19437,2024. URLhttps://doi.
org/10.48550/arXiv.2412.19437.
DeepSeek-AI. Deepseek-v2: Astrong,economical,andefficientmixture-of-expertslanguage
model. CoRR,abs/2405.04434,2024. URLhttps://doi.org/10.48550/arXiv.2405.04
434.
DeepSeek-AI. Fire-flyerfilesystem,2025. URLhttps://github.com/deepseek-ai/3FS.
DeepSeek-AI. Deepseek-r1incentivizesreasoninginllmsthroughreinforcementlearning. Nat.,
645(8081):633–638,2025. URLhttps://doi.org/10.1038/s41586-025-09422-z.
DeepSeek-AI. Deepseek-v3.2: Pushingthefrontierofopenlargelanguagemodels,2025. URL
https://arxiv.org/abs/2512.02556.
X.Deng,J.Da,E.Pan,Y.Y.He,C.Ide,K.Garg,N.Lauffer,A.Park,N.Pasari,C.Rane,K.Sampath,
M.Krishnan, S.Kundurthy, S.Hendryx,Z.Wang,V.Bharadwaj,J.Holm,R.Aluri,C.B.C.
Zhang,N.Jacobson,B.Liu,andB.Kenstler. Swe-benchpro: Canaiagentssolvelong-horizon
softwareengineeringtasks?,2025. URLhttps://arxiv.org/abs/2509.16941.
H.Ding,Z.Wang,G.Paolini,V.Kumar,A.Deoras,D.Roth,andS.Soatto. Fewertruncations
improvelanguagemodeling. arXivpreprintarXiv:2404.10830,2024.
X.Dong,Y.Fu,S.Diao,W.Byeon,Z.CHEN,A.S.Mahabaleshwarkar,S.-Y.Liu,M.V.keirsbilck,
M.-H.Chen,Y.Suhara,Y.C.Lin,J.Kautz,andP.Molchanov. Hymba: Ahybrid-headarchi-
tectureforsmalllanguagemodels. InTheThirteenthInternationalConferenceonLearning
Representations,2025. URLhttps://openreview.net/forum?id=A1ztozypga.
X.Du,Y.Yao,K.Ma,B.Wang,T.Zheng,K.Zhu,M.Liu,Y.Liang,X.Jin,Z.Wei,etal. Supergpqa:
Scalingllmevaluationacross285graduatedisciplines. arXivpreprintarXiv:2502.14739,2025.
D.Dua,Y.Wang,P.Dasigi,G.Stanovsky,S.Singh,andM.Gardner. DROP:Areadingcompre-
hensionbenchmarkrequiringdiscretereasoningoverparagraphs.InJ.Burstein,C.Doran,and
T.Solorio,editors,Proceedingsofthe2019ConferenceoftheNorthAmericanChapterofthe
Association for Computational Linguistics: Human Language Technologies, NAACL-HLT
2019,Minneapolis,MN,USA,June2-7,2019,Volume1(LongandShortPapers),pages2368–
2378.AssociationforComputationalLinguistics, 2019. doi: 10.18653/V1/N19-1246. URL
https://doi.org/10.18653/v1/n19-1246.
X.Gao,M.Dong,X.Miao,W.Du,C.Yu,andH.Chen. Erofs: acompression-friendlyreadonly
file system for resource-scarce devices. In Proceedings of the 2019 USENIX Conference on
UsenixAnnualTechnicalConference,USENIXATC’19,page149–162,USA,2019.USENIX
Association. ISBN9781939133038.
47

---

A.P.Gema, J.O.J.Leang, G.Hong, A.Devoto, A.C.M.Mancino, R.Saxena, X.He, Y.Zhao,
X. Du, M. R. G. Madani, C. Barale, R. McHardy, J. Harris, J. Kaddour, E. van Krieken, and
P.Minervini. Arewedonewithmmlu? CoRR,abs/2406.04127,2024. URLhttps://doi.or
g/10.48550/arXiv.2406.04127.
F. Gloeckle, B. Y. Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve. Better & faster large
language models via multi-token prediction. In Forty-first International Conference on
MachineLearning,ICML2024,Vienna,Austria,July21-27,2024.OpenReview.net,2024. URL
https://openreview.net/forum?id=pEWAcejiU2.
L. Haas, G. Yona, G. D’Antonio, S. Goldshtein, and D. Das. Simpleqa verified: A reliable
factuality benchmark to measure parametric knowledge. arXiv preprint arXiv:2509.07968,
2025.
Y. He, S. Li, J. Liu, Y. Tan, W. Wang, H. Huang, X. Bu, H. Guo, C. Hu, B. Zheng, et al. Chi-
nese simpleqa: A chinese factuality evaluation for large language models. arXiv preprint
arXiv:2411.07140,2024.
D.Hendrycks,C.Burns,S.Basart,A.Zou,M.Mazeika,D.Song,andJ.Steinhardt. Measuring
massivemultitasklanguageunderstanding. arXivpreprintarXiv:2009.03300,2020.
D.Hendrycks,C.Burns,S.Kadavath,A.Arora,S.Basart,E.Tang,D.Song,andJ.Steinhardt.Mea-
suringmathematicalproblemsolvingwiththemathdataset. arXivpreprintarXiv:2103.03874,
2021.
Y.Huang,Y.Bai,Z.Zhu,J.Zhang,J.Zhang,T.Su,J.Liu,C.Lv,Y.Zhang,J.Lei,etal. C-Eval: A
multi-levelmulti-disciplinechineseevaluationsuiteforfoundationmodels. arXivpreprint
arXiv:2305.08322,2023.
D.HupkesandN.Bogoychev. Multiloko: amultilinguallocalknowledgebenchmarkforllms
spanning31languages. CoRR,abs/2504.10356,2025. doi: 10.48550/ARXIV.2504.10356. URL
https://doi.org/10.48550/arXiv.2504.10356.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko.
Quantizationandtrainingofneuralnetworksforefficientinteger-arithmetic-onlyinference.
InProceedingsoftheIEEEConferenceonComputerVisionandPatternRecognition(CVPR),
June2018.
N.Jain,K.Han,A.Gu,W.-D.Li,F.Yan,T.Zhang,S.Wang,A.Solar-Lezama,K.Sen,andI.Stoica.
Livecodebench: Holisticandcontaminationfreeevaluationoflargelanguagemodelsforcode.
arXivpreprintarXiv:2403.07974,2024.
K.Jordan,Y.Jin,V.Boza,J.You,F.Cesista,L.Newhouse,andJ.Bernstein. Muon: Anoptimizer
forhiddenlayersinneuralnetworks. Citedon,page10,2024.
M.Joshi,E.Choi,D.Weld,andL.Zettlemoyer. TriviaQA:Alargescaledistantlysupervisedchal-
lengedatasetforreadingcomprehension.InR.BarzilayandM.-Y.Kan,editors,Proceedingsof
the55thAnnualMeetingoftheAssociationforComputationalLinguistics(Volume1: Long
Papers), pages 1601–1611, Vancouver, Canada, July 2017. Association for Computational
Linguistics. doi: 10.18653/v1/P17-1147. URLhttps://aclanthology.org/P17-1147.
H. Li, Y. Yuan, R. Du, K. Ma, L. Liu, and W. Hsu. DADI: Block-Level image service for agile
andelasticapplicationdeployment. In2020USENIXAnnualTechnicalConference(USENIX
ATC 20), pages 727–740. USENIX Association, July 2020. ISBN 978-1-939133-14-4. URL
https://www.usenix.org/conference/atc20/presentation/li-huiba.
48

---

H.Li,Y.Zhang,F.Koto,Y.Yang,H.Zhao,Y.Gong,N.Duan,andT.Baldwin. CMMLU:Measur-
ingmassivemultitasklanguageunderstandinginChinese. arXivpreprintarXiv:2306.09212,
2023.
J.Li, W.Zhao, J.Zhao, W.Zeng, H.Wu, X.Wang, R.Ge, Y.Cao, Y.Huang, W.Liu, etal. The
tooldecathlon: Benchmarkinglanguageagentsfordiverse,realistic,andlong-horizontask
execution. arXivpreprintarXiv:2510.25726,2025.
Y. Li, F. Wei, C. Zhang, and H. Zhang. EAGLE: speculative sampling requires rethinking
featureuncertainty. InForty-firstInternationalConferenceonMachineLearning,ICML2024,
Vienna,Austria,July21-27,2024.OpenReview.net,2024. URLhttps://openreview.net
/forum?id=1NdN7eXyb4.
J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y. Du, Y. Qin, W. Xu, E. Lu, J. Yan, Y. Chen, H. Zheng,
Y.Liu,S.Liu,B.Yin,W.He,H.Zhu,Y.Wang,J.Wang,M.Dong,Z.Zhang,Y.Kang,H.Zhang,
X. Xu, Y. Zhang, Y. Wu, X. Zhou, and Z. Yang. Muon is scalable for LLM training. CoRR,
abs/2502.16982,2025. URLhttps://doi.org/10.48550/arXiv.2502.16982.
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. arXiv preprint
arXiv:1711.05101,2017.
K.LuandT.M.Lab. On-policydistillation. ThinkingMachinesLab: Connectionism,2025. doi:
10.64434/tml.20251026. https://thinkingmachines.ai/blog/on-policy-distillation.
Z.Lu,C.Li,Y.Shi,W.Shen,M.Yan,andF.Huang. Corpusqa: A10milliontokenbenchmark
forcorpus-levelanalysisandreasoning. arXivpreprintarXiv:2601.14952,2026.
T. Luong, D. Hwang, H. H. Nguyen, G. Ghiasi, Y. Chervonyi, I. Seo, J. Kim, G. Bingham,
J. Lee, S. Mishra, A. Zhai, C. H. Hu, H. Michalewski, J. Kim, J. Ahn, J. Bae, X. Song, T. H.
Trinh, Q. V. Le, and J. Jung. Towards robust mathematical reasoning. In Proceedings of
the 2025 Conference on Empirical Methods in Natural Language Processing, 2025. URL
https://aclanthology.org/2025.emnlp-main.1794/.
M.A.Merrill,A.G.Shaw,N.Carlini,B.Li,H.Raj,I.Bercovich,L.Shi,J.Y.Shin,T.Walshe,E.K.
Buchanan,etal. Terminal-bench: Benchmarkingagentsonhard,realistictasksincommand
lineinterfaces. arXivpreprintarXiv:2601.11868,2026.
MiniMax. Meetminimax-m2,2025. URLhttps://github.com/MiniMax-AI/MiniMax-M2.
L. d. Moura and S. Ullrich. The lean 4 theorem prover and programming language. In
InternationalConferenceonAutomatedDeduction,pages625–635.Springer,2021.
Y.Nesterov. Amethodofsolvingaconvexprogrammingproblemwithconvergencerate𝑂(1/𝑘2).
SovietMathematicsDoklady,27:372–376,1983.
NVIDIACorporation. cublasdocumentation,2024. URLhttps://docs.nvidia.com/cuda
/cublas/. Version12.4.Accessed: 2024-09-16.
OpenAI. Multilingual massive multitask language understanding (mmmlu), 2024a. URL
https://huggingface.co/datasets/openai/MMMLU.
OpenAI. Openai mrcr: Long context multiple needle in a haystack benchmark, 2024b. URL
https://huggingface.co/datasets/openai/mrcr.
49

---

OpenAI. Learningtoreasonwithllms,2024c. URLhttps://openai.com/index/learnin
g-to-reason-with-llms.
OpenAI. IntroducingSimpleQA,2024d. URLhttps://openai.com/index/introducing
-simpleqa/.
OpenAI. Introducing SWE-bench verified we’re releasing a human-validated subset of swe-
benchthatmore,2024e. URLhttps://openai.com/index/introducing-swe-bench-v
erified/.
OpenAI. gpt-oss-120b&gpt-oss-20bmodelcard. CoRR,abs/2508.10925,2025. doi: 10.48550/A
RXIV.2508.10925. URLhttps://doi.org/10.48550/arXiv.2508.10925.
M.Osama,D.Merrill,C.Cecka,M.Garland,andJ.D.Owens. Stream-k: Work-centricparallel
decompositionfordensematrix-matrixmultiplicationonthegpu. InProceedingsofthe28th
ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming,
pages429–431,2023.
T. Patwardhan, R. Dias, E. Proehl, G. Kim, M. Wang, O. Watkins, S. P. Fishman, M. Aljubeh,
P. Thacker, L. Fauconnet, et al. Gdpval: Evaluating ai model performance on real-world
economicallyvaluabletasks. arXivpreprintarXiv:2510.04374,2025.
L.Phan,A.Gatti,Z.Han,N.Li,J.Hu,H.Zhang,C.B.C.Zhang,M.Shaaban,J.Ling,S.Shi,etal.
Humanity’slastexam. arXivpreprintarXiv:2501.14249,2025.
W.Qi,Y.Yan,Y.Gong,D.Liu,N.Duan,J.Chen,R.Zhang,andM.Zhou. Prophetnet: Predicting
future n-gram for sequence-to-sequence pre-training. In T. Cohn, Y. He, and Y. Liu, edi-
tors,FindingsoftheAssociationforComputationalLinguistics: EMNLP2020,OnlineEvent,
16-20November2020,volumeEMNLP2020ofFindingsofACL,pages2401–2410.Associa-
tionforComputationalLinguistics,2020. URLhttps://doi.org/10.18653/v1/2020.f
indings-emnlp.217.
Qwen. Qwen3technicalreport. CoRR,abs/2505.09388,2025. doi: 10.48550/ARXIV.2505.09388.
URLhttps://doi.org/10.48550/arXiv.2505.09388.
S.Rajbhandari,J.Rasley,O.Ruwase,andY.He.Zero: Memoryoptimizationstowardtrainingtril-
lionparametermodels. InSC20: InternationalConferenceforHighPerformanceComputing,
Networking,StorageandAnalysis,pages1–16.IEEE,2020.
J.K.Reed,Z.DeVito,H.He,A.Ussery,andJ.Ansel. Torch.fx: Practicalprogramcaptureand
transformationfordeeplearninginpython,2022. URLhttps://arxiv.org/abs/2112.0
8429.
D.Rein,B.L.Hou,A.C.Stickland,J.Petty,R.Y.Pang,J.Dirani,J.Michael,andS.R.Bowman.
GPQA:Agraduate-levelgoogle-proofq&abenchmark. arXivpreprintarXiv:2311.12022,2023.
G. T. M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard,
B.Shahriari,A.Ram’e,J.Ferret,P.Liu,P.D.Tafti,A.Friesen,M.Casbon,S.Ramos,R.Kumar,
C.L.Lan,S.Jerome,A.Tsitsulin,N.Vieillard,P.Stan´czyk,S.Girgin,N.Momchev,M.Hoff-
man, S. Thakoor, J.-B. Grill, B. Neyshabur, A. Walton, A. Severyn, A. Parrish, A. Ahmad,
A. Hutchison, A. Abdagic, A. Carl, A. Shen, A. Brock, A. Coenen, A. Laforge, A. Paterson,
B.Bastian,B.Piot,B.Wu,B.Royal,C.Chen,C.Kumar,C.Perry,C.A.Welty,C.A.Choquette-
Choo,D.Sinopalnikov,D.Weinberger,D.Vijaykumar,D.Rogozi’nska,D.Herbison,E.Bandy,
E.Wang,E.Noland,E.Moreira,E.Senter,E.Eltyshev,F.Visin,G.Rasskin,G.Wei,G.Cameron,
50

---

G. Martins, H. Hashemi, H. Klimczak-Pluci’nska, H. Batra, H. Dhand, I. Nardini, J. Mein,
J.Zhou,J.Svensson,J.Stanway,J.Chan,J.Zhou,J.Carrasqueira,J.Iljazi,J.Becker,J.Fernan-
dez,J.R.vanAmersfoort,J.Gordon,J.Lipschultz,J.Newlan,J.Ji,K.Mohamed,K.Badola,
K.Black,K.Millican,K.McDonell,K.Nguyen,K.Sodhia,K.Greene,L.L.Sjoesund,L.Usui,
L.Sifre,L.Heuermann,L.ciaLago,L.McNealus,L.B.Soares,L.Kilpatrick,L.Dixon,L.L.B.
Martins, M. Reid, M. Singh, M. Iverson, M. Gorner, M. Velloso, M. Wirth, M. Davidow,
M.Miller,M.Rahtz,M.Watson,M.Risdal,M.Kazemi,M.Moynihan,M.Zhang,M.Kahng,
M.Park,M.Rahman,M.Khatwani,N.Dao,N.shadBardoliwalla,N.Devanathan,N.Dumai,
N.Chauhan,O.Wahltinez,P.Botarda,P.Barnes,P.Barham,P.Michel,P.chongJin,P.Georgiev,
P.Culliton,P.Kuppala,R.Comanescu,R.Merhej,R.Jana,R.A.Rokni,R.Agarwal,R.Mullins,
S. Saadat, S. M. M. Carthy, S. Perrin, S. M. R. Arnold, S. bastian Krause, S. Dai, S. Garg,
S.Sheth,S.Ronstrom,S.Chan,T.Jordan,T.Yu,T.Eccles,T.Hennigan,T.Kociský,T.Doshi,
V.Jain,V.Yadav,V.Meshram,V.Dharmadhikari,W.Barkley,W.Wei,W.Ye,W.Han,W.Kwon,
X.Xu,Z.Shen,Z.Gong,Z.Wei,V.Cotruta,P.Kirk,A.Rao,M.Giang,L.Peran,T.Warkentin,
E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, D. Sculley, J. Banks, A. Dragan, S. Petrov,
O.Vinyals,J.Dean,D.Hassabis,K.Kavukcuoglu,C.Farabet,E.Buchatskaya,S.Borgeaud,
N. Fiedel, A. Joulin, K. Kenealy, R. Dadashi, and A. Andreev. Gemma 2: Improving open
languagemodelsatapracticalsize. arXivpreprintarXiv:2408.00118,2024.
S. Roller, S. Sukhbaatar, A. Szlam, and J. Weston. Hash layers for large sparse models. In
M.Ranzato,A.Beygelzimer,Y.N.Dauphin,P.Liang,andJ.W.Vaughan,editors,Advances
in Neural Information Processing Systems 34: Annual Conference on Neural Information
Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 17555–17566,
2021. URL https://proceedings.neurips.cc/paper/2021/hash/92bf5e6240737
e0326ea59846a83e076-Abstract.html.
B.D.Rouhani,R.Zhao,A.More,M.Hall,A.Khodamoradi,S.Deng,D.Choudhary,M.Cornea,
E.Dellinger,K.Denolf,S.Dusan,V.Elango,M.Golub,A.Heinecke,P.James-Roxby,D.Jani,
G.Kolhe,M.Langhammer,A.Li,L.Melnick,M.Mesmakhosroshahi,A.Rodriguez,M.Schulte,
R. Shafipour, L. Shao, M. Siu, P. Dubey, P. Micikevicius, M. Naumov, C. Verrilli, R. Wittig,
D.Burger,andE.Chung. Microscalingdataformatsfordeeplearning,2023.
K.Sakaguchi,R.L.Bras,C.Bhagavatula,andY.Choi. Winogrande: Anadversarialwinograd
schemachallengeatscale,2019.
Z.Shao,Y.Luo,C.Lu,Z.Z.Ren,J.Hu,T.Ye,Z.Gou,S.Ma,andX.Zhang. Deepseekmath-v2:
Towardsself-verifiablemathematicalreasoning,2025. URLhttps://arxiv.org/abs/25
11.22570.
N.Shazeer. Fasttransformerdecoding: Onewrite-headisallyouneed. CoRR,abs/1911.02150,
2019. URLhttp://arxiv.org/abs/1911.02150.
N.Shazeer. Gluvariantsimprovetransformer. arXivpreprintarXiv:2002.05202,2020.
F.Shi,M.Suzgun,M.Freitag,X.Wang,S.Srivats,S.Vosoughi,H.W.Chung,Y.Tay,S.Ruder,
D.Zhou,D.Das,andJ.Wei. Languagemodelsaremultilingualchain-of-thoughtreasoners.
In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali,
Rwanda,May1-5,2023.OpenReview.net,2023. URLhttps://openreview.net/forum?i
d=fR3wGCk-IXp.
J.Su,M.Ahmed,Y.Lu,S.Pan,W.Bo,andY.Liu. Roformer: Enhancedtransformerwithrotary
positionembedding. Neurocomputing,568:127063,2024.
51

---

M.Suzgun,N.Scales,N.Schärli,S.Gehrmann,Y.Tay,H.W.Chung,A.Chowdhery,Q.V.Le,
E.H.Chi,D.Zhou,etal. Challengingbig-benchtasksandwhetherchain-of-thoughtcansolve
them. arXivpreprintarXiv:2210.09261,2022.
G. Tsoukalas, J. Lee, J. Jennings, J. Xin, M. Ding, M. Jennings, A. Thakur, and S. Chaudhuri.
Putnambench: Evaluatingneuraltheorem-proversontheputnammathematicalcompetition,
2024. URLhttps://arxiv.org/abs/2407.11214.
A.Vaswani,N.Shazeer,N.Parmar,J.Uszkoreit,L.Jones,A.N.Gomez,Ł.Kaiser,andI.Polo-
sukhin. Attention is all you need. Advances in neural information processing systems, 30,
2017.
L.Wang,H.Gao,C.Zhao,X.Sun,andD.Dai. Auxiliary-loss-freeloadbalancingstrategyfor
mixture-of-experts. CoRR,abs/2408.15664,2024a. URLhttps://doi.org/10.48550/arX
iv.2408.15664.
L.Wang, Y.Cheng, Y.Shi, Z.Mo, Z.Tang, W.Xie, T.Wu, L.Ma, Y.Xia, J.Xue, etal. Tilelang:
Bridge programmability and performance in modern neural kernels. In The Fourteenth
InternationalConferenceonLearningRepresentations,2026.
Y.Wang,X.Ma,G.Zhang,Y.Ni,A.Chandra,S.Guo,W.Ren,A.Arulraj,X.He,Z.Jiang,T.Li,
M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen. Mmlu-pro: A more robust and
challengingmulti-tasklanguageunderstandingbenchmark. CoRR,abs/2406.01574,2024b.
URLhttps://doi.org/10.48550/arXiv.2406.01574.
J.Wei,Z.Sun,S.Papay,S.McKinney,J.Han,I.Fulford,H.W.Chung,A.T.Passos,W.Fedus,
andA.Glaese. Browsecomp: Asimpleyetchallengingbenchmarkforbrowsingagents. arXiv
preprintarXiv:2504.12516,2025.
T.Wei,J.Luan,W.Liu,S.Dong,andB.Wang. Cmath: Canyourlanguagemodelpasschinese
elementaryschoolmathtest?,2023.
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis. Efficient streaming language models with
attentionsinks. InTheTwelfthInternationalConferenceonLearningRepresentations,ICLR
2024,Vienna,Austria,May7-11,2024.OpenReview.net,2024. URLhttps://openreview
.net/forum?id=NG7sS51zVF.
Z.Xie,Y.Wei,H.Cao,C.Zhao,C.Deng,J.Li,D.Dai,H.Gao,J.Chang,K.Yu,L.Zhao,S.Zhou,
Z.Xu,Z.Zhang,W.Zeng,S.Hu,Y.Wang,J.Yuan,L.Wang,andW.Liang. mhc: Manifold-
constrainedhyper-connections,2026. URLhttps://arxiv.org/abs/2512.24880.
L.Xu,H.Hu,X.Zhang,L.Li,C.Cao,Y.Li,Y.Xu,K.Sun,D.Yu,C.Yu,Y.Tian,Q.Dong,W.Liu,
B.Shi,Y.Cui,J.Li,J.Zeng,R.Wang,W.Xie,Y.Li,Y.Patterson,Z.Tian,Y.Zhang,H.Zhou,
S.Liu,Z.Zhao,Q.Zhao,C.Yue,X.Zhang,Z.Yang,K.Richardson,andZ.Lan. CLUE:Achi-
neselanguageunderstandingevaluationbenchmark. InD.Scott,N.Bel,andC.Zong,editors,
Proceedings of the 28th International Conference on Computational Linguistics, COLING
2020, Barcelona, Spain (Online), December 8-13, 2020, pages 4762–4772. International Com-
mitteeonComputationalLinguistics,2020. doi: 10.18653/V1/2020.COLING-MAIN.419. URL
https://doi.org/10.18653/v1/2020.coling-main.419.
J.Yang,K.Lieret,C.E.Jimenez,A.Wettig,K.Khandpur,Y.Zhang,B.Hui,O.Press,L.Schmidt,
andD.Yang. Swe-smith: Scalingdataforsoftwareengineeringagents, 2025. URLhttps:
//arxiv.org/abs/2504.21798.
52

---

R.Zellers,A.Holtzman,Y.Bisk,A.Farhadi,andY.Choi. HellaSwag: Canamachinereallyfinish
yoursentence? InA.Korhonen,D.R.Traum,andL.Màrquez,editors,Proceedingsofthe57th
ConferenceoftheAssociationforComputationalLinguistics,ACL2019,Florence,Italy,July
28-August2,2019,Volume1: LongPapers,pages4791–4800.AssociationforComputational
Linguistics,2019. doi: 10.18653/v1/p19-1472. URLhttps://doi.org/10.18653/v1/p1
9-1472.
C. Zhang, K. Du, S. Liu, W. Kwon, X. Mo, Y. Wang, X. Liu, K. You, Z. Li, M. Long,
J. Zhai, J. Gonzalez, and I. Stoica. Jenga: Effective memory management for serving llm
with heterogeneity. In Proceedings of the ACM SIGOPS 31st Symposium on Operating
Systems Principles, SOSP ’25, page 446–461, New York, NY, USA, 2025a. Association for
Computing Machinery. ISBN 9798400718700. doi: 10.1145/3731569.3764823. URL
https://doi.org/10.1145/3731569.3764823.
S.Zhang,N.Zheng,H.Lin,Z.Jiang,W.Bao,C.Jiang,Q.Hou,W.Cui,S.Zheng,L.-W.Chang,
Q. Chen, and X. Liu. Comet: Fine-grained computation-communication overlapping for
mixture-of-experts. 2025b. URLhttps://arxiv.org/abs/2502.19811.
C.Zhao,L.Zhao,J.Li,Z.Xu,andC.Xu. Deepgemm: cleanandefficientfp8gemmkernelswith
fine-grainedscaling. https://github.com/deepseek-ai/DeepGEMM,2025.
W.Zhong,R.Cui,Y.Guo,Y.Liang,S.Lu,Y.Wang,A.Saied,W.Chen,andN.Duan. AGIEval: A
human-centricbenchmarkforevaluatingfoundationmodels. CoRR,abs/2304.06364,2023.
doi: 10.48550/arXiv.2304.06364. URLhttps://doi.org/10.48550/arXiv.2304.06364.
D. Zhu, H. Huang, Z. Huang, Y. Zeng, Y. Mao, B. Wu, Q. Min, and X. Zhou. Hyper-
connections. InTheThirteenthInternationalConferenceonLearningRepresentations,ICLR
2025,Singapore,April24-28,2025.OpenReview.net,2025. URLhttps://openreview.net
/forum?id=9FqARW7dwB.
X.Zhu,D.Cheng,H.Li,K.Zhang,E.Hua,X.Lv,N.Ding,Z.Lin,Z.Zheng,andB.Zhou. How
tosynthesizetextdatawithoutmodelcollapse? arXivpreprintarXiv:2412.14689,2024.
T.Y.Zhuo,M.C.Vu,J.Chim,H.Hu,W.Yu,R.Widyasari,I.N.B.Yusuf,H.Zhan,J.He,I.Paul,
S. Brunner, C. Gong, J. Hoang, A. R. Zebaze, X. Hong, W. Li, J. Kaddour, M. Xu, Z. Zhang,
P. Yadav, and et al. Bigcodebench: Benchmarking code generation with diverse function
calls and complex instructions. In The Thirteenth International Conference on Learning
Representations,ICLR2025,Singapore,April24-28,2025.OpenReview.net,2025. URLhttp
s://openreview.net/forum?id=YrycTjllL0.
53

---

Appendix
A. Author List and Acknowledgment
A.1. AuthorList
Authorsarelistedalphabeticallybytheirfirstname. Namesmarkedwith*denoteindividuals
whohavedepartedfromourteam.
Research & Engineering: Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang*, Bingzheng Xu,
BochaoWu,BoweiZhang,ChaofanLin,ChenDong,ChengdaLu,ChenggangZhao,Chengqi
Deng,ChenhaoXu,ChenzeShao,ChongRuan*,ConnerSun,DamaiDai,DayaGuo*,Dejian
Yang, Deli Chen, Donghao Li, Erhang Li, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong
Dai,GuangboHao,GuantingChen,GuoaiCao,GuolaiMeng,GuoweiLi,HanYu,HanZhang,
HanweiXu,HaoLi,HaofenLiang,HaolingZhang,HaomingLuo,HaoranWei*,HaotianYuan,
HaoweiZhang*,HaowenLuo,HaoyuChen,HaozheJi,HonghuiDing,HongxuanTang,Huanqi
Cao, Huazuo Gao, Hui Qu, Hui Zeng, J. Yang, J.Q. Zhu, Jia Yu, Jialiang Huang, Jiasheng Ye,
JiashiLi,JiaxinXu,JiewenHu,JinYan,JingchangChen,JingliZhou,JingtingXiang,Jingyang
Yuan,JingyuanCheng,JinhuaZhu,JipingYu,JosephSun,JunRan*,JunguangJiang,JunjieQiu,
JunlongLi*,JunxiaoSong,KaiDong,KaigeGao,KangGuan,KexingZhou,KezhaoHuang*,
Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Li Zhang, Liang Zhao, Lihua Guo, Lingxiao
Luo, Linwang Ma, Litong Wang, Liyu Cai, Liyue Zhang, Longhao Chen, M.S. Di, M.Y Xu,
MaxMei,MingchuanZhang,MinghuaZhang,MinghuiTang,MingxuZhou,PanpanHuang,
PeixinCong,PeiyiWang,QianchengWang,QihaoZhu,QingyangLi,QinyuChen,QiushiDu,
QiweiJiang,RuiTian,RuifanXu,RuijieLu,RuilingXu,RuiqiGe,RuisongZhang,RuizhePan,
RunjiWang,RunqianChen,RunqiuYin,RunxinXu,RuomengShen,RuoyuZhang,S.H.Liu,
ShanghaoLu,ShangyanZhou,ShanhuangChen,ShaofeiCai,ShaohengNie,ShaoyuanChen,
ShengdingHu,ShengyuLiu,ShiqiangHu,ShirongMa,ShiyuWang,ShuipingYu,Shunfeng
Zhou,ShutingPan,ShuyingYu,SongyangZhou,TaoNi,TaoYun,TianJin,TianPei,TianYe,
TianleLin,TianranJi,TianyiCui,TianyuanYue,TingtingYu,TunWang,W.Zhang,Wangding
Zeng,WeilinZhao,WenLiu,WenfengLiang,WenjiePang,WenjingLuo,WenjingYao,Wenjun
Gao,WenkaiYang,WenlveHuang,WentaoZhang,WentingMa,XiGao,XiangHe,Xiangwen
Wang,XiaoBi,XiaodongLiu,XiaohanWang,XiaokangChen,XiaokangZhang,XiaotaoNie,
XinCheng,XinLiu,XinXie,XingchaoLiu,XingchenLiu,XingkaiYu,XingyouLi,XinyuYang,
XuChen,XuanyuWang,XuechengSu,XuhengLin,XuweiFu,Y.C.Yan,Y.Q.Wang*,Y.W.Ma,
YanfengLuo,YangZhang,YanhongXu,YanruMa,YanwenHuang,YaoLi,YaoLi,YaoZhao,
YaofengSun,YaohuiWang,YiQian,YiYu,YichaoZhang,YifanDing,YifanShi,YijiaWu,Yiliang
Xiong,YingHe,YingZhou,YingjiaLuo,YinminZhong,YishiPiao,YisongWang,YixiangZhang,
YixiaoChen,YixuanTan,YixuanWei,YiyangMa,YiyuanLiu,YonglunYang,YongqiangGuo,
YongtongWu,YuWu,YuanCheng,YuanOu,YuanfanXu,YuanhaoLi,YuduanWang,Yuhan
Wu,YuhaoMeng,YuhengZou,YuKunLi,YunfanXiong,YupengChen,YuqianCao,Yuqian
Wang,YushunZhang,YutongLin,YuxianGu,YuxiangLuo,YuxiangYou,YuxuanLiu,Yuxuan
Zhou,YuyangZhou,YuzhenHuang,Z.F.Wu,ZehaoWang,ZehuaZhao,ZehuiRen,Zhangli
Sha,ZheFu,ZheanXu,ZhendaXie,ZhengyanZhang,ZhewenHao,ZhibinGou,ZhichengMa,
ZhigangYan,ZhihongShao,ZhixianHuang,ZhixuanChen,ZhiyuWu,ZhizhouRen,Zhuoshu
Li,ZhupingZhang,ZianXu,ZihaoWang,ZihuiGu,ZijiaZhu,ZilinLi,ZipengZhang*,Ziwei
Xie,ZiyiGao,ZizhengPan,ZongqingYao.
Business&Compliance: ChenchenLing,ChengyuHou,DongjieJi,FangWei,HengqingZhang,
JiaLuo,JiaSong,JialuCai,JianLiang,JiangtingZhou,JieyuYang,JinChen,JingziZhou,Junmin
Zheng,LeyiXia,LinyanZhu,MiaojunWang,MingmingLi,MinminHan,NingWang,Panpan
54

---

Wang, PengZhang, RuyiChen, ShangmianSun, ShaoqingWu, W.L.Xiao, WeiAn, Wenqing
Hou, Xianzu Wang, Xiaowen Sun, Xiaoxiang Wang, Xinyu Zhang, Xueyin Chen, Yao Xu, Yi
Shao,YilingMa,YingTang,YuehanYang,YuerXu,YukunZha,YupingLin,YutingYan,Zekai
Zhang,ZheJu,ZherenGao,ZhongyuWu,ZihuaQu,ZiyiWan.
A.2. Acknowledgment
WewouldliketothankDollyDengandothertestersfortheirvaluablesuggestionsandfeedback
regardingthecapabilitiesofDeepSeek-V4seriesmodels.
B. Evaluation Details
Table9 | AgenticSearchvs. RetrievalAugmentedSearchforDeepSeek-V4-Pro.
Difficulty Category # AgentWin RAGWin Tie Agent% RAG% Tie%
ObjectiveQ&A(客观问答) 196 110 43 43 56.1 21.9 21.9
Easy
SubjectiveQ&A(主观问答) 321 198 56 67 61.7 17.4 20.9
ObjectiveQ&A(客观问答) 168 102 33 33 60.7 19.6 19.6
Hard
SubjectiveQ&A(主观问答) 184 126 27 31 68.5 14.7 16.8
Total(总计) 869 536 159 174 61.7 18.3 20.0
Table 10 | Cost Comparison:Agentic Search vs. Retrieval Augmented Search (Mean) for
DeepSeek-V4-Pro. MostofthetoolcallsareparallelforAgenticSearch.
Version ToolCalls Prefill(tokens) Output(tokens)
V4AgenticSearch 16.2 13649 1526
V4RetrievalAugmentedSearch — 10453 1308
Table 11 | Comparative Evaluation of DeepSeek-V4-Pro and DeepSeek-V3.2 on Search Q&A
Tasks.
InternalEvaluation(内部综合评估)
Category Subcategory # V4win V3.2win tie V4% V3.2% tie%
Single-valueSearch(单值信息查找) 95 36 10 49 37.9 10.5 51.6
Objective
EntitySearch(实体信息查找) 99 24 7 68 24.2 7.1 68.7
Q&A
EnumerativeSearch(枚举型信息查找) 95 19 8 68 20.0 8.4 71.6
(客观问答)
Subtotal(小计) 289 79 25 185 27.3 8.7 64.0
CausalAnalysis(原因分析) 100 28 5 67 28.0 5.0 67.0
Comparison(对比) 96 28 20 48 29.2 20.8 50.0
AdviceSeeking(寻求建议) 92 23 8 61 25.0 8.7 66.3
Subjective
Recommendation(推荐) 95 26 19 50 27.4 20.0 52.6
Q&A
Planning&Strategy(攻略计划) 92 32 11 49 34.8 12.0 53.3
(主观问答)
Opinion&Evaluation(评价看法) 96 30 8 58 31.2 8.3 60.4
TrendAnalysis(趋势分析) 96 23 3 70 24.0 3.1 72.9
Subtotal(小计) 667 190 74 403 28.5 11.1 60.4
TOTAL(总计) 956 269 99 588 28.1 10.4 61.5
55

---

Figure14 | Exampleoutputofataskthatrequirescomparingtworegularinvestmentstrategies
fortheNASDAQ.
Figure15 | Exampleoutputofataskwhichrequiresresearching2020-2025NobelSciencePrizes
andgeneratingananalyticalPDFreport.
56

---

Table12 | ComparativeAnalysisofDeepSeek-V4-ProandGemini-3.1-ProinChineseFunctional
Writing.
InternalEvaluation(内部综合评估)
Category Subcategory # DSwin Gemwin Tie DS% Gem% Tie%
Report(报告) 527 350 162 15 66.41 30.74 2.85
Proposal(方案策划) 291 181 103 7 62.20 35.40 2.41
Education(教育培训) 159 100 56 3 62.89 35.22 1.89
Email&Letter(邮件书信) 146 107 37 2 73.29 25.34 1.37
Business
Writing Notice(通知公告) 72 43 24 5 59.72 33.33 6.94
(办公文本) Professional(专业文本) 63 34 27 2 53.97 42.86 3.17
Recruitment(招聘求职) 42 27 15 0 64.29 35.71 0.00
Technical(技术文本) 29 22 7 0 75.86 24.14 0.00
Review(介绍评价) 20 15 5 0 75.00 25.00 0.00
Subtotal(小计) 1349 879 436 34 65.16 32.32 2.52
SocialMedia(社交媒体文案) 267 156 101 10 58.43 37.83 3.75
AdCopy(广告商品文案) 214 109 98 7 50.93 45.79 3.27
Long-formContent(内容平台长文) 99 71 25 3 71.72 25.25 3.03
Media NewsReport(新闻报道) 51 27 22 2 52.94 43.14 3.92
Writing Advertorial(营销软文) 17 12 4 1 70.59 23.53 5.88
(媒体文本) Headline(标题) 11 7 4 0 63.64 36.36 0.00
NarrationScript(口播文案) 4 2 1 1 50.00 25.00 25.00
Comment(评论) 3 2 1 0 66.67 33.33 0.00
Subtotal(小计) 666 386 256 24 57.96 38.44 3.60
Congratulatory(祝贺文本) 101 54 41 6 53.47 40.59 5.94
Communication(沟通回复) 100 71 26 3 71.00 26.00 3.00
Everyday
Reflection(心得感想) 90 68 17 5 75.56 18.89 5.56
Writing
Review(介绍评价) 55 44 9 2 80.00 16.36 3.64
(生活文本)
Comment(评论) 44 34 8 2 77.27 18.18 4.55
Subtotal(小计) 390 271 101 18 69.49 25.90 4.62
Speech(发言稿) 226 135 85 6 59.73 37.61 2.65
NarrationScript(口播文案) 51 25 23 3 49.02 45.10 5.88
Oral
SalesScript(话术) 31 22 6 3 70.97 19.35 9.68
Writing
(口头文本) Dialogue(对话文本) 10 4 6 0 40.00 60.00 0.00
Congratulatory(祝贺文本) 1 1 0 0 100.00 0.00 0.00
Subtotal(小计) 319 187 120 12 58.62 37.62 3.76
AdministrativeDoc(事务文书) 117 60 53 4 51.28 45.30 3.42
PersonalDoc(个人文书) 73 45 27 1 61.64 36.99 1.37
Official GovernmentDoc(行政公文) 34 19 14 1 55.88 41.18 2.94
Document
(公文文本) Speech(发言稿) 3 1 2 0 33.33 66.67 0.00
EssayWriting(申论写作) 3 1 1 1 33.33 33.33 33.33
Subtotal(小计) 230 126 97 7 54.78 42.17 3.04
ResearchPaper(学术论文) 104 67 32 5 64.42 30.77 4.81
Academic Coursework(课程作业) 90 53 35 2 58.89 38.89 2.22
Writing AcademicSupport(学术辅助) 15 11 3 1 73.33 20.00 6.67
(学术文本) ScienceOutreach(专业科普) 7 6 1 0 85.71 14.29 0.00
Subtotal(小计) 216 137 71 8 63.43 32.87 3.70
Total(总计) 3170 1986 1081 103 62.65 34.10 3.25
57

---

Table13 | ComparativeAnalysisofDeepSeek-V4-ProandGemini-3.1-ProinChineseCreative
Writing.
InstructionFollowing(指令遵循) WritingQuality(写作质量)
Subcategory(文体) # DS Gem Tie DS% Gem% Tie% DS Gem Tie DS% Gem% Tie%
Fiction(小说故事) 836 504 323 5 60.58 38.82 0.60 672 157 3 80.77 18.87 0.36
GeneralFiction(泛小说故事) 662 368 290 3 55.67 43.87 0.45 467 194 0 70.65 29.35 0.00
FanFiction(同人文) 410 253 150 3 62.32 36.95 0.74 338 67 1 83.25 16.50 0.25
GeneralFanFic.(泛同人文) 202 111 90 1 54.95 44.55 0.50 161 40 1 79.70 19.80 0.50
Narrative(记叙文) 171 115 54 2 67.25 31.58 1.17 141 30 0 82.46 17.54 0.00
GeneralProse(泛散文) 124 83 40 1 66.94 32.26 0.81 88 36 0 70.97 29.03 0.00
Prose(散文) 112 74 38 0 66.07 33.93 0.00 92 20 0 82.14 17.86 0.00
WritingStyle(文笔) 112 81 31 0 72.32 27.68 0.00 86 26 0 76.79 23.21 0.00
ClassicalPoetry(古诗文) 48 24 24 0 50.00 50.00 0.00 39 9 0 81.25 18.75 0.00
ModernPoetry(现代诗) 43 23 20 0 53.49 46.51 0.00 32 11 0 74.42 25.58 0.00
Lyrics(歌词) 30 8 22 0 26.67 73.33 0.00 16 14 0 53.33 46.67 0.00
LiteraryAppreciation(赏析) 27 20 7 0 74.07 25.93 0.00 18 9 0 66.67 33.33 0.00
GeneralArgument.(泛议论文) 24 15 9 0 62.50 37.50 0.00 17 7 0 70.83 29.17 0.00
GeneralNarrative(泛记叙文) 23 11 12 0 47.83 52.17 0.00 15 8 0 65.22 34.78 0.00
GeneralClassical(泛古文诗歌) 9 5 4 0 55.56 44.44 0.00 5 4 0 55.56 44.44 0.00
CreativeWriting(创意写作) 6 2 4 0 33.33 66.67 0.00 4 2 0 66.67 33.33 0.00
Argumentative(议论文) 5 5 0 0 100.00 0.00 0.00 5 0 0 100.00 0.00 0.00
GeneralMod.Poetry(泛现代诗) 2 1 1 0 50.00 50.00 0.00 2 0 0 100.00 0.00 0.00
Total(总计) 2837 1703 1119 15 60.03 39.44 0.53 2198 634 5 77.48 22.35 0.18
Table14 | DeepSeek-V4-Provs. Claude-Opus-4.5onComplexInstructionFollowingandMulti-
TurnWriting.
InternalEvaluation(内部综合评估)
Category # DS Opus Tie DS% Opus% Tie%
ComplexInst. Following(复杂指令跟随) 49 23 26 0 46.9% 53.1% 0.0%
Multi-TurnWriting(多轮写作) 147 67 76 4 45.6% 51.7% 2.7%
Total(总计) 196 90 102 4 45.9% 52.0% 2.0%
58

---

