Several specialized models can read the same document without being able to reuse one another’s cached computation directly. Their inputs match, but their weights differ. A coding model and a checking model may therefore spend GPU time rebuilding internal state for context that has already been processed elsewhere.

DroidSpeak investigates how much of that work can be avoided[1]. The University of Chicago and Microsoft collaboration profiles fine-tuned models derived from the same foundation, recomputes selected layers, and reuses other layers’ KV state. Its objective is to reduce prefill cost while retaining acceptable measured task quality.

The qualifier matters. This is not ordinary exact caching for the same model, and matching architecture alone does not guarantee that reuse will work. The paper’s limitations explicitly retain a same-foundation scope and acknowledge that even some pairs within that scope may not benefit. A deployment must validate a particular sender, receiver, and task distribution.

Why identical text does not make an identical cache

During prefill, a transformer processes the input and creates key and value tensors used by subsequent decoding. Those tensors depend on the weights and intermediate representations, not just the input text. Fine-tuning can change the representation needed to preserve a specialized model’s behavior.

If the receiver simply accepts every KV layer from the sender, it can skip substantial computation. However, it is then decoding against state that its own prefill would not necessarily have produced. In the paper’s experiments, this direct substitution can cause large quality losses. A fast cache transfer does not establish a valid replacement for the intended computation.

That is a different problem from finding repeated text across requests to one fixed model. It is also different from caching documents before retrieval or moving exact state between replicas. DroidSpeak permits an approximation across weight versions and evaluates the consequences using task metrics.

For an operator, the distinction changes the acceptance test. Cache integrity, correct ordering, and successful transfer remain necessary, but they are not sufficient. The resulting answer must also meet the application’s quality requirements under the changed state. A transport checksum cannot verify that semantic property.

Layer sensitivity as an empirical opportunity

The study tests eight model pairs over datasets covering question answering, summarization, and code completion. It examines what happens when particular layers reuse sender state while other layers are recomputed. The effect is not uniform: some layers are much more sensitive to the change than others.

Across the examined pairs, the authors identify an average of roughly 11% of layers as critical under their sensitivity test. That is an empirical observation, not a theorem about transformers. It also does not mean the final runtime recomputes exactly 11% of every model, because efficient and accurate recomputation has additional dependencies.

The identity of sensitive layers is comparatively stable across the studied inputs for a given pair. This motivates profiling before deployment rather than making an expensive layer-selection decision for every request. The system can then choose among previously measured configurations that trade recomputation against quality.

However, sensitivity depends on the model pair and the direction of reuse. A sender with one specialization and a receiver with another do not necessarily behave like the reversed pair. A useful cache relationship therefore needs more metadata than a shared base-model name; it needs the validated configuration for the actual direction and versions being served.

The activation required to restart computation

Recomputing an internal layer requires its input activation for the complete context. Reused KV state alone does not supply that input. Running every earlier layer again would recreate much of the prefill work the system is trying to avoid.

DroidSpeak instead obtains an activation tensor from the sender at the point where reuse changes to recomputation. The paper calls this E cache. It allows the receiver to start the selected region without recomputing the full prefix of the network, but the activation itself is also an approximation to what the receiver would have produced.

Choosing several scattered sensitive layers can therefore be expensive in two ways. It requires multiple transition activations to be stored and transferred, and it introduces multiple points at which sender-state differences enter recomputation. Selecting only the individually most sensitive layers does not automatically produce the best complete configuration.

The design consequently profiles contiguous groups. Recomputing some less-sensitive layers between critical ones can avoid another transition and its associated state error. The optimization target is the quality and latency of the whole selected region, not the smallest count of layers that looked important in isolation.

A receiver reuses selected KV layers but needs a sender activation at the transition into a contiguous recomputation region. Recomputing the intervening layers can avoid additional transitions. The layer positions are illustrative, not a recommended configuration for a named model. Original figure created for this article.

Profiling makes quality a configuration constraint

Offline profiling tests recomputation choices on representative data and records the best observed quality for each amount of work. The runtime can select a configuration that satisfies a quality target while reducing computation. That target must be specified by the application rather than inferred from the word negligible.

The main experimental procedure uses 50 HotpotQA contexts for profiling and applies the configurations to the evaluated test datasets. A 5% quality-loss threshold appears in the profiling discussion, while the online-serving comparisons select configurations within a 1% drop for the shown pairs. Those are experimental settings, not proof of identical output or a guaranteed limit on every request.

The metrics also have different meanings. F1 measures answer overlap for the question-answering evaluation; Rouge-L measures a text-overlap property for summarization; code similarity measures closeness to reference code. None alone establishes factual safety, functional correctness on all programs, or unchanged behavior for a particular customer.

Profiling has a measurable cost. For the reported 32-layer glue_sst2/conllpp pair, per-layer granularity takes 3.6 hours, two-layer groups take one hour, and three-layer groups take 0.375 hours. Coarser search reduces work, but the observed preservation of the quality tradeoff in those tests does not eliminate the need to check new pairs.

Measured profiling times for the paper’s 32-layer glue_sst2/conllpp case: 3.6 hours at one-layer granularity, 1 hour for two-layer groups, and 0.375 hours for three-layer groups. These are offline profiling costs for one evaluated pair, not inference latencies. Original figure created for this article.

Transfer order determines how much computation is saved

Remote state must arrive before the receiver can use it. Loading every cached layer before starting recomputation serializes two substantial activities and can erase part of the computational benefit. Loading only the KV layers that will be reused is better, but still leaves avoidable waiting if computation begins only after all transfers finish.

The key dependency is narrower: once the transition activation arrives, the receiver can start recomputing its selected region. Other reusable KV layers can continue transferring while that computation proceeds. The implementation uses separate CUDA streams for transfer and computation to expose this overlap.

This does not make network cost disappear. It determines which part can be hidden behind useful computation. If the network is slow or the reusable state is large, transfer can remain limiting. If the network is very fast, additional overlap has less value because there is less transfer delay left to hide.

The paper’s runtime adapts recomputation using system load and a latency objective. Automatically changing the policy to reflect varying network bandwidth is listed as future work. We should therefore distinguish the measured benefit of pipelining at different bandwidths from an implemented controller that continuously optimizes for all network conditions.

The memory cost of keeping reusable state

Sender state must exist when another model needs it. Alongside KV caches, the system retains the transition activation required by the selected configuration. One example for Llama 3.1 8B at 10K tokens uses about 1.2 GB for all KV layers and 0.08 GB for one E layer, an additional amount of roughly 6% in that case.

The example is useful, but it is not a universal memory ratio. Tensor dimensions, precision, attention structure, context length, and the number of retained transitions can change the cost. Storing caches for many contexts or model versions also multiplies the capacity that must remain available long enough to produce reuse.

This creates a lifecycle question that raw prefill timing does not answer. If a context is evicted before another model consumes it, the system pays storage and profiling overhead without obtaining the expected reuse. If it is retained too long, it can displace state needed by active requests. The cache policy should be evaluated against actual cross-model reuse intervals.

Versioning and access control are also necessary deployment concerns. A cache identifier must distinguish the model state and validated direction of reuse, while authorization must prevent one tenant’s private context from becoming available to another merely because text hashes or model names match. These are requirements derived from the design, not claims about a completed security evaluation in the paper.

What the experiments establish

The hardware consists of two A100 virtual machines, each with eight 80 GB GPUs, connected through 200 Gbps InfiniBand. The evaluated 70B models use four-bit AWQ quantization to fit on one GPU. That precision choice belongs to the experiment and should remain visible when comparing quality or memory requirements with another deployment.

Across the reported pairs and datasets, prefill speedups range from 1.7× to 3.1×, averaging 2.1× in the paper’s summary. Prefill here includes loading the relevant remote state as well as computation. The comparison is with full receiver-model prefill, while direct full-cache reuse offers a different and often much worse quality tradeoff.

The online experiment uses model replicas on the two nodes, round-robin routing, and Poisson request arrivals. At latency objectives chosen to avoid the full-prefill system’s sharp queueing increase, the reported request-throughput improvement is 2–4×. That is not network bandwidth and not a guarantee that every agent program finishes four times sooner.

A separate coding-agent case uses a coder and tester with HumanEval tasks. At a configuration preserving the reported pass@1 score of 52.5, the study reports a 2.7× TTFT improvement and lower end-to-end completion delay. This provides an application-level example, but does not expand the result into a general guarantee for arbitrary agent teams or coding workloads.

Conditions that require a new validation

The most immediate limitation is model or data drift. Updating a fine-tuned model changes the pair that was profiled. Serving substantially different inputs can change which layers are sensitive and how much quality is lost. The paper explicitly identifies periodic reprofiling as a possible response rather than presenting it as a fully solved online mechanism.

The experiments primarily validate pairs. A chain in which one model reuses approximated state and another then reuses that result raises a different question about accumulated error. A successful pairwise test does not establish that an arbitrary multi-hop chain retains the same quality, and optimizing reuse among many models remains future work.

For a pilot, we would begin with a frequently reused context class and a stable, versioned model pair. Compare against full prefill using task-specific outcomes, not just aggregate text metrics. Retain a full-computation fallback and measure the cost of profiling, cache retention, and remote transfer alongside the saved GPU execution.

The useful systems idea is that related models need not repeat every layer of context processing if a validated approximation can preserve the required behavior. The benefit comes from combining that empirical opportunity with the right activation boundary and transfer schedule. It should be deployed as a quality-controlled optimization, not treated as a new universal meaning of cache compatibility.

Sources and rights

This independently written analysis is based on the final NSDI 2026 paper, including its limitations and appendices. Measurements are attributed to that evaluation; deployment considerations are editorial analysis. Original-paper copyright remains with the authors under USENIX publication terms, © 2026. All graphics are newly created from the described mechanism or reported data, without reproducing original figure arrangements.