Long-context training에서는 같은 token 수가 같은 작업량을 뜻하지 않습니다. Packed sequence 하나가 long document 하나를 담을 수도 있고 여러 short document를 담을 수도 있습니다. Attention mask는 document 경계를 넘어선 attention을 막으므로 long document의 뒤쪽 token이 짧은 document 묶음의 같은 절대 위치 token보다 더 많은 key를 봅니다. Fixed sequence length는 memory 형상만 맞출 뿐 arithmetic intensity는 맞추지 못합니다.
Meta 내부의 405B training trace는 이 비용을 8,000개 H100에서 보여 줍니다. 128K context window에서 가장 느린 GPU의 attention computation latency는 다른 GPU보다 1.44배 길었습니다[1]. Synchronous collective와 pipeline dependency는 이 지연을 바깥 계층으로 전파합니다. 따라서 large training system의 유효 capacity는 token 총량이나 nominal FLOPs가 아니라 가장 무거운 document 조합이 결정합니다.
WLB-LLM은 4D training flow의 두 경계를 바꿉니다. Pipeline parallelism(PP)에서는 micro-batch마다 다른 token length를 허용해 predicted total latency를 맞춥니다. Context parallelism(CP)에서는 packed document를 각각 shard할 수 있게 하되, 균형 이점이 attention kernel 효율 감소보다 클 때만 finer layout을 선택합니다. Input composition을 immutable loader output이 아니라 runtime scheduling variable로 다루는 구조입니다.
Equal-token partition 뒤에 남는 두 imbalance
Framework는 데이터, pipeline, context, tensor 병렬화를 결합합니다. Data 병렬 worker가 global batch를 받고 pipeline stage는 micro-batch sequence를 실행합니다. Context parallelism은 long sequence를 GPU 사이에 나누며 tensor parallelism은 model operation을 분할합니다. 각 boundary가 같은 token 수를 보더라도 synchronization rule은 서로 다릅니다.
PP level에서는 long document 하나가 든 micro-batch가 여러 short document를 packing한 micro-batch보다 attention work가 큽니다. 모든 pipeline stage는 같은 micro-batch composition을 처리하고 가장 느린 micro-batch가 전체 dependency 연결 경로를 통과합니다. 그 latency에 첫 stage에 남은 forward와 backward work가 더해지므로 packing에서 시작한 skew가 pipeline에서 증폭됩니다.
CP level의 일반 layout은 sequence 전체를 CP worker 수의 두 배 chunk로 자르고 symmetric pair를 배정합니다. 하나의 uninterrupted document라면 early region과 late attention region을 짝지어 balance합니다. 여러 document를 packing하면 attention이 다시 시작하는 위치가 달라집니다. 같은 token 수라도 document tail이 많이 든 chunk의 query-key work가 더 커집니다. CP worker 안의 tensor 병렬 worker는 동일한 gathered chunk를 처리하므로 TP level에서 이 차이가 사라지지 않습니다.

Variable length로 확보한 balancing room
Fixed-length repacking에는 명확한 한계가 있습니다. Document 하나가 context window를 이미 채웠다면 quadratic attention work를 맞출 short document 묶음은 token limit을 넘어야 합니다. 여러 global batch로 packing window를 늘리면 수학적으로 더 균형 잡히지만 더 많은 training example의 실행 순서를 바꿉니다. 논문의 550M pretraining 실험에서는 window가 커질수록 final training loss가 증가했으므로 scheduler freedom에도 model-quality cost가 있음을 보여 줍니다.
WLB-LLM은 memory에서 정한 upper bound 안에서 micro-batch마다 다른 length를 허용합니다. Objective는 document length로 attention latency를 예측하고 GEMM, elementwise operation, collective latency를 더합니다. 여러 short document는 masked attention work가 낮다면 nominal context length보다 긴 sequence가 될 수 있습니다. 이때 늘어난 non-attention work가 outlier의 total latency에 가까워지므로 unrelated document를 큰 reorder window로 옮기지 않아도 됩니다.
Runtime algorithm은 pending document를 length 순으로 보고 predicted work가 가장 작은 micro-batch에 greedily 배정합니다. Memory bound를 넘으면 현재 length가 가장 짧은 micro-batch를 시도하고, 그래도 들어가지 않는 document는 다음 iteration에 남깁니다. Attention과 other-operation latency function은 offline profile에서 얻습니다. 이 model은 모든 환경에 통하는 token formula가 아닙니다. GPU, kernel, precision, communication mapping이 바뀌면 profile도 다시 만들어야 합니다.
Global shuffle 대신 제한적으로 지연하는 outlier
Global batch 하나에는 모든 micro-batch에 비슷한 long document를 하나씩 줄 만큼 표본이 없을 수 있습니다. WLB-LLM은 unusually long document를 length band별 FIFO queue에 넣습니다. Queue가 micro-batch 수만큼 쌓여야 각 micro-batch에 similar outlier를 하나씩 release합니다. Threshold interval을 좁히면 work는 더 잘 맞지만 entry가 모이는 시간은 길어집니다.
저자들은 training corpus 일부에서 threshold를 조정해 imbalance와 token delay를 함께 봤습니다. Outlier queue 두 개를 사용하면 reported imbalance metric은 1.05였고 batch당 packing 시간은 20ms로 step latency의 0.65% 미만이었습니다. Fixed-length integer-programming solver는 비슷한 balance를 얻을 수 있지만 global batch 네 개를 함께 풀 때 batch당 25초를 넘었습니다.
Token 지연은 평균 0.5 training iteration이었습니다. WLB-LLM의 loss curve는 global batch 하나 안에서 fixed-length packing한 결과와 같은 trend를 보였고 더 많은 batch를 repacking하면 loss가 늘었습니다. 이 근거는 평가한 pretraining setup에는 유효하지만 모든 optimizer, curriculum, rare-domain mixture, exact-resumption policy가 같은 delay를 허용한다는 뜻은 아닙니다. 운영 환경에서는 global average뿐 아니라 domain별 delay distribution도 추적해야 합니다.
Work balance와 tile efficiency의 tradeoff
CP boundary에서 WLB-LLM은 각 document를 CP group size의 두 배로 나누고 worker마다 symmetric piece를 줍니다. 나머지 token은 round-robin으로 분배해 explicit padding 없이 total token count도 맞춥니다. 각 worker가 모든 document의 early region과 late region을 비슷하게 받으므로 token count와 attention work가 함께 균형을 이룹니다.
Finer sharding은 kernel을 느리게 만들 수 있습니다. FlashAttention 계열 구현은 tile 단위로 계산하므로 short query fragment도 더 긴 fragment와 같은 tile을 차지해 computation을 낭비합니다. Hopper GPU는 query region이 충분히 크면 Tensor Memory Accelerator multicast와 L2를 통해 key-value load를 reuse할 수 있습니다. Document 하나를 작은 piece로 나누면 reuse가 줄고 achieved TFLOPs가 떨어집니다. Perfect work balance가 minimum step time을 보장하지 않는 이유입니다.
WLB-LLM은 micro-batch마다 candidate 두 개를 평가합니다. Per-sequence와 per-document sharding의 query 및 key-value 형상을 계산하고 kernel tile에 맞춰 반올림한 뒤 FLOPs를 offline profile의 achieved TFLOPs로 나눕니다. Predicted maximum CP-worker latency가 작은 layout을 선택합니다. 7B-128K breakdown에서 per-document sharding만 적용하면 1.02배였고 adaptive selection은 CP contribution을 1.05배로 높였습니다. 더 큰 효과는 PP packing과 delay에서 나온 1.28배였으며 둘을 결합하면 1.33배였습니다.
Context length에서 커지고 model size에서 줄어든 gain
평가는 node 32개에서 수행했습니다. Node마다 NVLink로 연결된 H100 SXM 80GB GPU 여덟 개를 사용하고 node 사이는 RoCE로 연결했습니다. Internal LLaMA-like model은 550M, 7B, 30B, 70B였으며 context window는 64K와 128K였습니다. 각 model scale은 다른 tensor, context, pipeline, 데이터 병렬 configuration을 사용했습니다. 가장 큰 평가 configuration은 GPU 256개였고 design 동기가 된 8,000-GPU trace와는 규모가 다릅니다.
Internal Plain-4D 대비 WLB-LLM은 평균 1.23배였고 Fixed-4D 대비 1.19배였습니다. Fixed-4D는 global batch 하나에서만 repacking하고 training run 전체에 CP sharding choice 하나를 사용해 평균 1.03배에 그쳤습니다. 7B model에서 context가 32K에서 160K로 늘면 speedup은 1.03배에서 1.40배로 커졌습니다. Main configuration에서 64K를 128K로 늘리면 평균은 1.15배에서 1.30배로 바뀌었습니다.
Larger model에서는 communication이 step time에서 더 큰 비율을 차지하므로 attention-work imbalance를 줄이는 최적화의 relative gain이 작았습니다. Longer context는 attention share와 outlier document probability를 함께 높여 addressable fraction을 키웁니다. Capacity plan에서 average 1.23배를 short-context, communication-bound, differently packed job에 그대로 적용하면 안 됩니다. 먼저 각 workload의 imbalance profile을 측정해야 합니다.
Training contract에 들어오는 scheduling metadata
WLB-LLM은 distributed-training balance에 tensor 형상뿐 아니라 document boundary가 필요함을 보여 줍니다. 실제 rollout은 loader, packing metadata, CP partitioner, attention kernel 전체에서 이 boundary를 보존해야 합니다. Restart 뒤 example order가 조용히 달라지지 않도록 checkpoint에는 deterministic queue state도 들어가야 합니다. Offline latency profile은 GPU architecture, kernel build, precision, 병렬 mapping과 함께 버전을 관리해야 합니다.
Mixture-of-Experts는 token이 expert로 routing된 뒤 다른 imbalance를 만듭니다. 논문은 packing과 sharding이 dropless expert routing의 gating decision을 바꾸지 않아 compatible하다고 설명합니다. 그렇더라도 expert hot spot 측정은 별도로 필요합니다. Input balance와 expert balance는 서로 다른 scheduler이며 worst case가 같은 step에서 겹칠 수 있습니다.
Procurement 판단을 단순히 GPU를 23% 적게 사도 된다는 결론으로 바꾸면 안 됩니다. WLB-LLM은 long-document composition이 synchronized idle time을 만들 때 installed GPU당 completed training work를 높입니다. 운영자는 PP와 CP boundary의 slowest-to-average attention ratio, step time 중 attention share, token-delay distribution, packing overhead, model-quality trajectory를 먼저 측정해야 합니다. Paper와 같은 failure mode가 확인되면 document-aware scheduling은 model mathematics를 바꾸지 않고 capacity를 회수할 수 있습니다. Communication이나 expert routing이 지배하면 다른 bottleneck이 투자 효과를 결정합니다.
출처와 저작권 안내
이 글은 Silicon & Systems가 작성한 편집 분석으로 architecture, evaluation, limitation을 우리 표현으로 다시 썼습니다. 원문의 문장, 표, 도판은 재수록하지 않았고 도판은 이 글을 위해 새로 만들었습니다. 전체 논문은 USENIX OSDI 2025 발표 페이지에서 확인할 수 있습니다. 저작권은 저자에게 있습니다. 2025.