A new systems paper, FluidPD (Prefill Decode), splits each request into two phases and lets a server borrow GPU time between them, or swap worker roles in place, when one side goes idle.
When someone sends a question to a chatbot, the model does two very different kinds of work in sequence. First it reads the entire prompt and prepares a response, a compute-heavy burst called prefill. Then it writes the answer one token at a time, a slower, memory-bound stream called decode. Real traffic is lopsided and bursty: a wave of long prompts can swamp the prefill side while decode workers sit half-idle, and a long stretch of short answers can leave prefill machines cold while decode stalls. The question for any large-scale serving system is what to do in that gap.
Most production stacks today split the two phases onto separate pools of GPUs and pin the split with a fixed ratio. That ratio is set when traffic looks roughly balanced, then watched from a distance. When one side falls behind, autoscalers eventually add capacity, but the reaction is slow, costs money, and does nothing about short bursts on the timescale where autoscalers cannot react. The work that an idle GPU could have done in that short window is left on the floor, and the user on the other end waits for a response that is already late.
A new systems paper called FluidPD proposes a different lever: instead of adding GPUs, let the serving system flex the split between prefill and decode in place, using the hardware that is already online. The paper introduces two mechanisms that target the two timescales of imbalance, both guided by lightweight signals that read resource pressure on each side before the latency targets start slipping.
For short bursts, FluidPD uses a mechanism called FluidToken. When decode workers have momentary slack, a bounded portion of the next prefill job is offloaded onto them. The work runs where there is spare capacity, and the result stitches back into the main prefill pipeline. The offload is capped, so a long decode stream is never crowded out by a flood of prompts.
For sustained imbalance, where one side of the fleet is consistently busier than the other for minutes or longer, FluidPD uses a second mechanism called FluidRole. A running worker is reassigned from prefill to decode or back, in place, without reloading the model weights or restarting the serving engine. The transition is cheap because the model is already in GPU memory; only the role, the scheduler binding, and the request queue change.
Both mechanisms are steered by pressure indices, simple per-side measurements of how close prefill and decode are to missing their latency budgets. The indices are cheap to compute and run ahead of the SLO violations they are meant to prevent, which is what makes the system able to react on the timescales where autoscalers cannot.
The paper reports that, across replays of production Azure LLM serving traces, FluidPD improves overall SLO attainment by up to 94.6 percentage points compared with a static configuration in SGLang, the open-source serving runtime the authors used as the baseline. That is a large number, and the limits of what it actually shows are part of the story.
The comparison is against a single baseline: the same serving stack with the prefill/decode ratio locked. FluidPD is not benchmarked head-to-head against other prefill/decode disaggregation systems such as DistServe, Mooncake, or Splitwise in the hydrated abstract, so the result measures how much an in-place elastic split helps relative to a static split, not how it ranks against the full field. The evaluation is also a trace replay, not a live multi-tenant deployment, so the 94.6 pp number describes how the system would have behaved on recorded Azure traffic, not how it behaves under real, multi-tenant load with competing workloads and noisy neighbors.
The mechanisms themselves are bounded by design. FluidToken offloads only a slice of prefill to decode workers, so it cannot turn a starved prefill side into a healthy one by itself when decode is also busy. FluidRole reassigns existing workers rather than adding new ones, so it cannot raise total fleet capacity. The contribution is a higher duty cycle on hardware operators already own, not a way to skip the next GPU purchase.
The public conversation about AI capacity is still framed as a GPU-count problem: how many accelerators, how many data centers, how many megawatts. The FluidPD paper is part of a smaller, less visible current of systems work that argues the next several percentage points of useful answers will come from running the GPUs operators already have at full duty cycle, and that the lever is software, not silicon.
For an operator evaluating the bet, the watch items are direct. A live deployment study, especially one that mixes FluidToken-style transient offload with FluidRole-style role reassignment under multi-tenant load, would tell the field whether trace-replay gains survive contact with production noise. A head-to-head against an autoscaler-equipped static split, and against at least one competing prefill/decode disaggregation system, would tell the field where FluidPD sits in the design space. Neither comparison is in the hydrated abstract, and neither is necessary to make the paper worth reading, but both are necessary before the result is treated as a deployment-ready recipe rather than a promising mechanism.