Kihagyás

Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

Authors: Ruoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, Mingxing Zhang
Affiliations: Moonshot AI, Tsinghua University
Date: 2026-04-22 (arXiv:2604.15039v1)
Link: https://arxiv.org/abs/2604.15039v1


Abstract Summary

Prefill-decode (PD) disaggregation has become standard for LLM serving, but deployment is limited by KVCache transfer bandwidth. Conventional dense-attention models generate huge KVCache traffic that keeps prefill and decode tightly coupled within a single high-bandwidth network domain. Recent hybrid-attention architectures reduce KVCache size substantially, making cross-cluster KVCache transport plausible. However, smaller KVCache alone is not sufficient — real workloads are bursty, request lengths are skewed, prefix caches are unevenly distributed, and inter-cluster bandwidth fluctuates.

PrfaaS (Prefill-as-a-Service) presents a cross-datacenter serving architecture that: - Selectively offloads long-context prefill to standalone, compute-dense prefill clusters - Transfers resulting KVCache over commodity Ethernet to local PD clusters for decode - Uses length-based threshold routing, bandwidth-aware scheduling, and cache-aware request placement

Results: With an internal 1T-parameter hybrid model, PrfaaS achieves 54% higher throughput than homogeneous PD and 32% higher than naive heterogeneous baselines, while consuming only modest cross-datacenter bandwidth (~13 Gbps of 100 Gbps link, 13%).


Key Concepts

Prefill-Decode (PD) Disaggregation

  • Prefill: compute-intensive (processes full input prompt)
  • Decode: memory-bandwidth-intensive (generates tokens one at a time)
  • Standard approach: separate these phases to optimize each independently
  • Problem: KVCache transfer between prefill and decode nodes is bandwidth bottleneck

KVCache Transfer Bottleneck

  • Dense attention: KVCache grows linearly with sequence length → huge bandwidth demand
  • 32K tokens on MiniMax-M2.5: ~60 Gbps per instance → requires RDMA-class fabric
  • This confines PD deployments to single datacenter with high-bandwidth interconnect

Hybrid Attention Architectures

  • Interleave small number of full-attention layers with many linear-complexity layers
  • Examples:
  • Kimi Delta Attention (KDA) — 3:1 linear:full ratio
  • Sliding Window Attention (SWA) — 5:1 ratio
  • Ring-2.5-1T — 7:1 ratio with Lightning attention
  • Result: 4-13× reduction in KVCache size vs dense models
  • KV throughput drops from ~60 Gbps to ~3-8 Gbps → commodity Ethernet viable

PrfaaS Architecture

Components: 1. PrfaaS clusters — standalone compute-dense clusters for long-context prefill 2. Local PD clusters — conventional PD serving for decode and short prefills 3. Network — intra-cluster RDMA + inter-cluster VPC/dedicated Ethernet 4. Hybrid prefix cache pool — separate KVCache groups for linear states and full-attention KVCache

Selective Offloading: - Length-based routing threshold t: only requests with l > t go to PrfaaS - Short requests stay on local PD path - With log-normal distribution (μ=9.90, σ=1.00, mean ~27K tokens), optimal t = 19.4K - ~50% of requests routed to PrfaaS

Bandwidth-Aware Scheduling: - Monitors egress utilization and queue depth - Adjusts routing threshold dynamically when congestion approaches - Layer-wise prefill pipelining to overlap generation with transmission - Multi-connection TCP to fully utilize available bandwidth


Throughput Model

System throughput limited by slowest stage:

Λmax = min(Θprfaas/p, Θpd-p/(1-p), Θpd-d)

Where: - Θprfaas = min(Nprfaas/Tprefill(llong), Bout/Skv(llong)) - Θpd-p = Np/Tprefill(lshort) - Θpd-d = Nd·BSmax/(Tdecode·Lout) - p = fraction routed to PrfaaS

Optimization variables: 1. Routing threshold t (determines p, llong, lshort) 2. PD-cluster prefill-to-decode ratio Np/Nd

Optimal configuration (1T hybrid model, 32 H200 PrfaaS + 64 H20 PD): - t = 19.4K tokens - Np = 3, Nd = 5 (within PD cluster) - Mean TTFT: 2.22s (vs 4.44s homogeneous) - P90 TTFT: 3.51s (vs 9.73s homogeneous) - Throughput: 3.24 req/s (vs 2.11 homogeneous = +54%)


Hardware Implications

Phase-Specialized Hardware

  • NVIDIA Rubin CPX: compute throughput for prefill
  • Groq LPU / Taalas HC1: extreme memory bandwidth for decode
  • Cross-datacenter KVCache removes need for these to share same RDMA fabric
  • Allows independent scaling of prefill and decode capacity

Bandwidth Requirements

  • Ring-2.5-1T (1T params, 7:1 hybrid ratio):
  • 32K tokens: ~170 Gbps cross-datacenter
  • 128K tokens: below 100 Gbps
  • 10,000 GPU datacenter: ~1.8 Tbps aggregate — within modern DC fabric capacity
  • Compare: dense models require 3.8 Tbps (MiniMax-M2.5) or 2.1 Tbps (Qwen3)

  • Mooncake (Moonshot AI): KVCache-centric disaggregated architecture — cited as foundation
  • Splitwise/DistServe: PD disaggregation from cost/power and goodput perspectives
  • Helix/Hetis/LLM-PQ: heterogeneous GPU/network optimization
  • CacheGen/CacheBlend/FusionRAG: KVCache compression and reuse
  • KIVI/KVQuant/H2O: KVCache quantization and importance-based eviction

Key Insights

  1. Model architecture + system design together: Hybrid attention reduces KVCache enough, but selective offloading + bandwidth-aware scheduling makes cross-datacenter practical
  2. Not all requests should be offloaded: Short requests stay local — only long prefills benefit from compute-dense hardware
  3. Commodity Ethernet sufficient: 13% of 100 Gbps link for realistic workloads — no RDMA required
  4. Independent scaling: Prefill and decode can scale separately across regions/clouds
  5. Next-gen hardware co-design: Rubin CPX for prefill, LPU for decode — each in optimal location

Quote

"KVCache-efficient model architectures are necessary but not sufficient for cross-datacenter heterogeneous serving. What makes the deployment practical is the combination of model-side KVCache reduction with system-side selective offloading and bandwidth-aware scheduling."

Vissza a tetejére