# Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

**Authors:** Ruoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, Mingxing Zhang  
**Affiliations:** Moonshot AI, Tsinghua University  
**Date:** 2026-04-22 (arXiv:2604.15039v1)  
**Link:** https://arxiv.org/abs/2604.15039v1

---

## Abstract Summary

Prefill-decode (PD) disaggregation has become standard for LLM serving, but deployment is limited by KVCache transfer bandwidth. Conventional dense-attention models generate huge KVCache traffic that keeps prefill and decode tightly coupled within a single high-bandwidth network domain. Recent hybrid-attention architectures reduce KVCache size substantially, making cross-cluster KVCache transport plausible. However, smaller KVCache alone is not sufficient — real workloads are bursty, request lengths are skewed, prefix caches are unevenly distributed, and inter-cluster bandwidth fluctuates.

**PrfaaS (Prefill-as-a-Service)** presents a cross-datacenter serving architecture that:
- Selectively offloads long-context prefill to standalone, compute-dense prefill clusters
- Transfers resulting KVCache over commodity Ethernet to local PD clusters for decode
- Uses length-based threshold routing, bandwidth-aware scheduling, and cache-aware request placement

**Results:** With an internal 1T-parameter hybrid model, PrfaaS achieves 54% higher throughput than homogeneous PD and 32% higher than naive heterogeneous baselines, while consuming only modest cross-datacenter bandwidth (~13 Gbps of 100 Gbps link, 13%).

---

## Key Concepts

### Prefill-Decode (PD) Disaggregation
- Prefill: compute-intensive (processes full input prompt)
- Decode: memory-bandwidth-intensive (generates tokens one at a time)
- Standard approach: separate these phases to optimize each independently
- Problem: KVCache transfer between prefill and decode nodes is bandwidth bottleneck

### KVCache Transfer Bottleneck
- Dense attention: KVCache grows linearly with sequence length → huge bandwidth demand
- 32K tokens on MiniMax-M2.5: ~60 Gbps per instance → requires RDMA-class fabric
- This confines PD deployments to single datacenter with high-bandwidth interconnect

### Hybrid Attention Architectures
- Interleave small number of full-attention layers with many linear-complexity layers
- Examples:
  * Kimi Delta Attention (KDA) — 3:1 linear:full ratio
  * Sliding Window Attention (SWA) — 5:1 ratio
  * Ring-2.5-1T — 7:1 ratio with Lightning attention
- Result: 4-13× reduction in KVCache size vs dense models
- KV throughput drops from ~60 Gbps to ~3-8 Gbps → commodity Ethernet viable

### PrfaaS Architecture

**Components:**
1. **PrfaaS clusters** — standalone compute-dense clusters for long-context prefill
2. **Local PD clusters** — conventional PD serving for decode and short prefills
3. **Network** — intra-cluster RDMA + inter-cluster VPC/dedicated Ethernet
4. **Hybrid prefix cache pool** — separate KVCache groups for linear states and full-attention KVCache

**Selective Offloading:**
- Length-based routing threshold t: only requests with l > t go to PrfaaS
- Short requests stay on local PD path
- With log-normal distribution (μ=9.90, σ=1.00, mean ~27K tokens), optimal t = 19.4K
- ~50% of requests routed to PrfaaS

**Bandwidth-Aware Scheduling:**
- Monitors egress utilization and queue depth
- Adjusts routing threshold dynamically when congestion approaches
- Layer-wise prefill pipelining to overlap generation with transmission
- Multi-connection TCP to fully utilize available bandwidth

---

## Throughput Model

System throughput limited by slowest stage:

Λmax = min(Θprfaas/p, Θpd-p/(1-p), Θpd-d)

Where:
- Θprfaas = min(Nprfaas/Tprefill(llong), Bout/Skv(llong))
- Θpd-p = Np/Tprefill(lshort)
- Θpd-d = Nd·BSmax/(Tdecode·Lout)
- p = fraction routed to PrfaaS

Optimization variables:
1. Routing threshold t (determines p, llong, lshort)
2. PD-cluster prefill-to-decode ratio Np/Nd

Optimal configuration (1T hybrid model, 32 H200 PrfaaS + 64 H20 PD):
- t = 19.4K tokens
- Np = 3, Nd = 5 (within PD cluster)
- Mean TTFT: 2.22s (vs 4.44s homogeneous)
- P90 TTFT: 3.51s (vs 9.73s homogeneous)
- Throughput: 3.24 req/s (vs 2.11 homogeneous = +54%)

---

## Hardware Implications

### Phase-Specialized Hardware
- NVIDIA Rubin CPX: compute throughput for prefill
- Groq LPU / Taalas HC1: extreme memory bandwidth for decode
- Cross-datacenter KVCache removes need for these to share same RDMA fabric
- Allows independent scaling of prefill and decode capacity

### Bandwidth Requirements
- Ring-2.5-1T (1T params, 7:1 hybrid ratio):
  * 32K tokens: ~170 Gbps cross-datacenter
  * 128K tokens: below 100 Gbps
  * 10,000 GPU datacenter: ~1.8 Tbps aggregate — within modern DC fabric capacity
- Compare: dense models require 3.8 Tbps (MiniMax-M2.5) or 2.1 Tbps (Qwen3)

---

## Related Work

- **Mooncake** (Moonshot AI): KVCache-centric disaggregated architecture — cited as foundation
- **Splitwise/DistServe**: PD disaggregation from cost/power and goodput perspectives
- **Helix/Hetis/LLM-PQ**: heterogeneous GPU/network optimization
- **CacheGen/CacheBlend/FusionRAG**: KVCache compression and reuse
- **KIVI/KVQuant/H2O**: KVCache quantization and importance-based eviction

---

## Key Insights

1. **Model architecture + system design together**: Hybrid attention reduces KVCache enough, but selective offloading + bandwidth-aware scheduling makes cross-datacenter practical
2. **Not all requests should be offloaded**: Short requests stay local — only long prefills benefit from compute-dense hardware
3. **Commodity Ethernet sufficient**: 13% of 100 Gbps link for realistic workloads — no RDMA required
4. **Independent scaling**: Prefill and decode can scale separately across regions/clouds
5. **Next-gen hardware co-design**: Rubin CPX for prefill, LPU for decode — each in optimal location

---

## Quote

> "KVCache-efficient model architectures are necessary but not sufficient for cross-datacenter heterogeneous serving. What makes the deployment practical is the combination of model-side KVCache reduction with system-side selective offloading and bandwidth-aware scheduling."