Model marketplaces have a bookkeeping problem that nobody advertises. Alibaba Cloud’s Model Studio serves thousands of hosted models behind one API, and the traffic follows the shape every marketplace knows: in the paper’s production sample, 94.1% of 779 models together received 1.35% of 167.6 million requests. Serving etiquette nonetheless demands that each of those cold models sit on a warm GPU, so 17.7% of a fleet of roughly 30,000 GPUs idled at under 0.2 requests per second apiece, while the hot models burst past their reservations on a different timescale. A SOSP 2025 paper from Alibaba and Peking University puts a number on the waste and then removes most of it[1]: Aegaeon, the pooling system now running in Model Studio’s beta, cut the GPUs behind a few dozen marketplace models from 1,192 to 213, an 82% reduction, and the claim survived three months of production traffic.

We summarize the argument in our own words below. The interesting part is not that pooling helps (everyone pools) but why every previous pooling approach hit a wall of two to three models per GPU, and what it took to break past it.

Why request-granularity scaling stalls at three

Two families of prior work frame the problem. Multiplexing packs several models onto one device and shares it spatially or temporally[3], but weights are the constraint: models in this workload average 25.1 GB, so an 80 GB card fits two, perhaps three. Auto-scaling systems in the serverless tradition[2] instead park weights in host memory or SSD and load models on demand, which lifts the memory ceiling, but they swap models only when a request finishes. The paper’s sharpest analytical move is a small theorem about that choice. If each model’s requests arrive as a Poisson process with rate λ and a request occupies the system for T seconds, the expected number of simultaneously active models is M·(1−e^(−λT)). Plug in the production numbers (λ=0.037, T=16.79 s) and 100 models keep 46.55 active on average, though the aggregate load is a mere 3.7 requests per second. Since a request-granularity system must keep every active model resident somewhere, its pooling ratio is bounded near 100/46.55, which is to say: below three models per GPU, no better than multiplexing. LLM requests are long, so “active” is a much weaker condition than “busy”, and head-of-line blocking does the rest: a new model’s first token waits for someone else’s entire generation.

Aegaeon’s answer is to make the scheduling quantum the token. Between any two tokens, the system may preempt the resident model, park its KV cache in host memory, load a different model, generate for a while under an explicit time budget, and switch back, provided every request still meets its deadlines. Service quality is defined accordingly, per token: a time-to-first-token target for the first token (10 s in production) and a time-between-tokens target for the rest (100 ms), with the observation that decoded output can be buffered, so a request that decodes at burst speed earns slack that the scheduler may spend serving somebody else’s model. Prefill and decoding get separate GPU pools and separate policies[5]: prefill instances run grouped first-come-first-served queues (requests for the same model batch into groups of up to eight, so one model switch amortizes across arrivals), while decoding instances run weighted round-robin over per-model batches, with each batch’s time quota computed from its token deadline slack and the switching costs of everything in the rotation.

The long tail, and the theorem that dooms request-level scaling. a, In the production sample, 94.1% of 779 marketplace models draw 1.35% of 167.6 M requests yet hold 17.7% of roughly 30,000 GPUs. b, With Poisson arrivals at rate λ and service time T, the expected active model count is M·(1−e^(−λT)); at the measured λ=0.037 and T=16.79 s, 100 sporadically invoked models keep about 47 active, so any system that switches models only at request boundaries pools fewer than three models per GPU. Token-level preemption schedules around this bound. Original figure created for this article.

Making a model switch cost less than a second

None of that scheduling matters if switching models costs tens of seconds, and out of the box it does: scaling a 13B vLLM instance down and back up takes up to 26.9 seconds, most of it nowhere near the weights. The paper’s accounting finds the time hiding in engine reinitialization (distributed executor setup, profiling passes, KV cache pinning), in garbage collection forced by fragmented device memory, and in serialized KV cache transfers. Aegaeon attacks each line item. Engine components are initialized once per instance and reused across models, with only weights and KV cache treated as model-specific; that alone removes over 80% of the latency. Device memory becomes a self-managed buffer with bump allocation, and the tensor library’s own allocator is bypassed by patching the parameter classes at load time, so back-to-back model loads never trigger a collection pass. Host memory holds a shared model cache and pinned staging buffers, giving pipelined, multi-threaded weight loads that come in under a second, and a prefetch stream can stage the next scheduled model alongside the running one, making roughly half of all switches effectively instantaneous. KV cache for models of different shapes lands in slab-allocated unified pools (fragmentation stays under 20%), and transfers synchronize through CUDA events at the granularity of individual requests, so a decoding instance starts generating for a request the moment its cache arrives rather than when the whole batch does. End to end, the switching sequence drops by 97%, to sub-second.

What a preemptive model switch costs, before and after. The unoptimized sequence (KV cache out, garbage collection, engine reinitialization, weight load, KV cache in) reaches 26.9 s for a 13B model. Component reuse removes the reinitialization, explicit memory management (bump-allocated device buffers, a shared host model cache, slab-allocated KV pools) removes the collection and accelerates loading, prefetch hides the load entirely on about half of switches, and per-request CUDA-event synchronization overlaps the cache traffic: 97% of the latency is gone, and switches complete in under a second. Original figure created for this article.

What the numbers say

On a 16-GPU H800 testbed against ServerlessLLM (with and without an oracle-assisted shortest-job-first variant) and MuxServe, Aegaeon sustains 2 to 2.5× the request arrival rate at equal service quality, or 1.5 to 9× the goodput, and serves up to 70 models on ten decoding GPUs, the seven-models-per-GPU headline. The production deployment is the more consequential result: a cross-region cluster of 213 H20 GPUs now carries twenty-eight models of 1.8 to 7B and nineteen of 32 to 72B, workloads that previously occupied 1,192 H20s, with average GPU utilization rising from the 13.3 to 33.9% range to 48.1% and no observed SLO violations over the monitored period. Note the choice of silicon: H20 is the export-compliant accelerator whose scarcity value in China is exactly why an 82% reduction in fleet size reads as a strategic number rather than an operations footnote.

The paper is also candid about where the approach thins out. Tighten the deadlines to a fifth of the defaults (2 s first token, 20 ms between tokens) and the slack that funds preemption disappears; in the strictest configuration static multiplexing wins, since it never pays a switching cost. Token-level pooling is a machine for monetizing loose SLOs on sporadic traffic, not a universal serving accelerator, and the authors position it accordingly: the hot, latency-critical models keep their dedicated instances, while the long tail consolidates.

Deployment results, with the regime marked. Production: 1,192 H20 GPUs reduced to 213 (82%) across 47 marketplace models, utilization up from 13.3-33.9% to 48.1%, three months in beta without observed SLO violations. Testbed: 2-2.5× higher sustainable arrival rates, 1.5-9× goodput against auto-scaling and multiplexing baselines, seven models per decoding GPU. The boundary: at 5× stricter deadlines the preemption slack vanishes and static multiplexing regains the lead. Original figure created for this article.

What we take from it

The KV-cache-centric school of serving design, which we met in Mooncake’s disaggregated architecture[6] and in the CXL-attached form in Beluga, treats cache placement as the organizing problem of inference. Aegaeon extends the same logic one level up: not only where a request’s cache lives, but which model owns the GPU at all, becomes a per-token decision backed by fast state movement. We believe the theorem is the durable contribution. E[m]=M·(1−e^(−λT)) is a one-line explanation for why serverless LLM serving disappointed everyone who tried it at request granularity, and it hands infrastructure teams a formula for predicting their own pooling ceiling from two measurable numbers. The 82% figure, meanwhile, deserves its caveats: it describes a curated set of 47 marketplace models under loose SLOs with redundancy on both sides of the comparison, not a fleet-wide rule. What it does establish is that the long tail of the model market, previously priced at one warm GPU per cold model, now has a market-clearing price closer to one-seventh of that, and that the missing ingredient was never exotic hardware but a scheduler willing to make decisions every hundred milliseconds.

Token deadlines turn pooling into admission control

Token-level preemption creates useful slack, but it does not create unlimited capacity. Every admitted request adds a first-token deadline, a stream of later-token deadlines, model weights that may need to become resident, and KV state that must survive preemption. A scheduler can satisfy each local choice and still admit a combination that has no feasible rotation. The service therefore needs an admission test before weighted round-robin begins.

That test should distinguish prefill and decode. Prefill has a large, irregular compute demand and can benefit from grouping requests for the same model. Decode has smaller repeated steps, but each active sequence carries state and a next-token deadline. A model with one long generation can occupy little arithmetic at any instant while keeping its weights and KV cache relevant for minutes. Request count and token rate are therefore insufficient on their own. The admission model needs resident bytes, expected decode duration, switch cost, and the deadline slack already promised to active users.

Buffering decoded tokens also needs a policy limit. Bursty generation can earn time for a model switch only while the client still receives tokens within the stated inter-token objective. Spending all accumulated slack on a cold model may improve aggregate utilization and produce a visibly uneven stream. Operators should report deadline attainment and the distribution of inter-token gaps, not only average throughput. The purpose of pooling is to sell spare intervals without turning interactive generation into a sequence of stalls.

State movement defines the real pooling ratio

The theorem explains why request-boundary scaling cannot pool a long tail effectively. Token-level scaling removes that particular bound, but a second bound remains in the data path. Weights move from host memory to GPU memory, and KV state moves in the opposite direction or between decode instances. The amount of state that can be transferred before the next deadline limits how many models can share one device.

This is why Aegaeon’s engine work is part of the architecture rather than an implementation detail. Reusing initialized components removes fixed latency. Explicit device-memory management removes allocator and garbage-collection surprises. Concurrent KV movement allows the transfer path to approach the schedule assumed by the controller. If any one of these regresses after an engine update, the feasible rotation shrinks even though the scheduling algorithm is unchanged.

A production dashboard should expose switch latency by model size, bytes moved for weights and KV cache, host-memory pressure, and deadline slack before and after each switch. It should also separate switches that were hidden by buffered tokens from switches that appeared in user-visible latency. A nominal seven-model pooling result is useful only with this movement ledger attached. A different PCIe topology, host DRAM bandwidth, or model-size distribution can produce a different safe ratio.

A marketplace needs fairness above utilization

The long tail is economically attractive because many models have sparse demand. It is also vulnerable to starvation. A few popular models can keep the decode pool busy, while a rarely used model repeatedly pays a cold-start cost. Weighted round-robin can protect token deadlines after admission, but the market still needs a rule for capacity reservation, queueing, and rejection across model owners.

One option is to assign each model a minimum service share and let unused share enter a common pool. Another is to price cold-state retention explicitly, so an owner can choose between keeping weights warm and accepting a longer first token. Either policy must account for KV state as well as weights. A model with few requests but unusually long contexts can consume more host memory than its request share suggests. Fairness should therefore be stated in service outcomes, such as admitted requests meeting TTFT and inter-token objectives, rather than equal GPU seconds.

Rollout should begin with the part of the catalog whose request rate is low enough to waste dedicated instances but high enough to measure repeatedly. The operator can compare dedicated service, request-boundary scaling, and token-level pooling under the same arrival trace. The comparison should include cold and warm starts, model-size mix, context and output lengths, host-memory use, and rejected requests. Only after deadline prediction is calibrated should the scheduler accept rarer models with little history.

We read Aegaeon as evidence that accelerator virtualization for LLMs must include time, state, and service promises. Memory multiplexing alone fits only the weights that are simultaneously resident. Serverless loading alone waits too long for requests to finish. Token-level scheduling joins the two, but it works because the runtime makes switching predictable enough for the scheduler to promise a next token. The durable product claim is not a fixed number of models per GPU. It is the number of distinct model services whose requests can meet declared token deadlines on a measured state-movement path.

Source and attribution

This article is an editorial summary prepared for Silicon & Systems. It restates the argument of the paper cited below in our own words. No text, figures or tables from the paper are reproduced here, and the figures on this page were created for this summary. The paper appeared at SOSP 2025; the authoritative version is in the Proceedings of the 31st ACM Symposium on Operating Systems Principles, (c) 2025 the authors, publication rights licensed to ACM.