추천 학습은 GPU job 하나가 아닙니다. Sparse feature 처리와 sample 구성은 큰 CPU 배포 환경을 사용하고 embedding 및 model update는 CPU나 GPU에서 실행되며, example은 과거 batch와 live event로 함께 들어옵니다. ByteDance에서 모델 하나의 하루 데이터는 5년 동안 20TB에서 160TB로 늘었고 하루 CPU virtual-core 수요는 150만 개에서 900만 개로 증가했습니다. 단일 모델의 전체 학습 이력은 20PB가 넘을 수 있습니다.
Primus는 이 증가가 scheduler, storage, learning mode마다 별도 stack을 만들지 않게 설계한 시스템입니다[1]. YARN과 Kubernetes, 여러 storage system, offline·online·mixed training 위에 하나의 선언형 계약을 제공합니다. 2025년 배치는 CPU virtual core 1천만 개 이상, GPU 수만 장, 추천 학습 데이터 7EB를 포함했습니다.
논문의 세 가지 통합 계층은 서로 다른 병목을 풉니다. Resource unification은 heterogeneous executor를 제어하고 개수나 크기를 바꿉니다. Data unification은 storage마다 framework code를 넣지 않고 source, transformation, task를 표현합니다. Training unification은 fresh stream과 historical batch를 섞어 오래된 분포를 잊지 않으면서 model update를 빠르게 유지합니다.
두 scheduler 위의 공통 resource object
Primus는 YARN이나 Kubernetes 세부를 training framework에 노출하는 대신 custom resource definition으로 job을 표현합니다. Training master가 job object를 감시하고 resource controller가 가용 pool의 executor를 할당합니다. CPU, memory, I/O, progress를 관측한 뒤 horizontal 또는 vertical scaling을 실행합니다. Training code는 아래 scheduler가 달라도 동일한 executor contract를 봅니다.
Horizontal scaling은 executor 수를 바꾸며 병렬 preprocessing처럼 다시 나눌 수 있는 작업에 맞습니다. Vertical scaling은 기존 executor의 CPU나 memory를 바꿔 process별 용량이 병목일 때 restart를 피합니다. Primus는 GPU 추가만을 유일한 elastic action으로 보지 않고 job과 cluster 상태에서 전략을 고릅니다.
한 운영 실험에서는 job당 CPU utilization이 약 50%에서 80%로 높아졌습니다. Cluster rollout에서 core당 training throughput은 30.26에서 35.44로 올라 17.1% 개선과 비용 절감으로 보고됐습니다. 다른 vertical-scaling 사례는 memory allocation을 고친 뒤 처리량이 초당 275에서 496 minibatch로 늘었습니다. 추천 algorithm이 빨라진 결과가 아니라 resource matching의 효과입니다.

Elasticity에는 안정적인 progress unit이 필요합니다. Executor 수를 바꾸면 데이터 ownership이나 collective communication이 흔들릴 수 있고 vertical change는 병목을 memory에서 I/O로 옮길 수 있습니다. Primus는 운영 telemetry와 보수적인 adjustment window를 사용합니다. 절감 수치는 관리된 배포 환경의 근거이며 모든 job을 계속 scaling해야 한다는 뜻은 아닙니다.
경로 문자열이 아닌 데이터 task graph
추천 sample은 distributed file, table, feature store, stream에 나뉘어 있을 수 있습니다. 하나의 path만 받는 training system은 preprocessing logic을 framework code에 넣고 새 storage format을 침습적 변경으로 만듭니다. Primus는 logical source definition, physical partition, executable task를 3계층 데이터 description으로 분리합니다.
Planner는 요청한 time range와 source mixture를 데이터 task graph로 전개합니다. Data executor가 sample을 읽고 변환한 뒤 RPC로 training executor에 제공합니다. Graph는 dependency와 parallelism을 표현하고 adapter는 storage-specific access를 담습니다. TensorFlow나 PyTorch가 모든 backend를 알지 않아도 batch와 stream input을 한 job에 넣을 수 있습니다.
두 source에서 20일 분량을 사용한 운영 입력에서 기존 serial task generation은 58분이 걸렸습니다. 분산 생성은 149초로 줄어 23배 빨랐고 planner thread 네 개에서는 42초까지 짧아졌습니다. 학습이 시작되기 전 control-plane bottleneck 때문에 accelerator가 쉬고 있었으며, model math를 바꾸지 않고 이를 제거했습니다.
Data execution은 여전히 skew에 민감합니다. Example이 많거나 storage가 느린 partition 하나가 step boundary를 잡을 수 있습니다. Primus는 task를 재분배하고 planning과 loading을 분리하지만 useful sample/s, storage traffic, executor CPU를 계속 측정해야 합니다. 빠른 graph generator가 균형 있는 graph execution을 자동 보장하지는 않습니다.
과거 분포를 버리지 않는 freshness
Pure online training은 stream에서 빠르게 갱신하지만 최근 분포에 과적합하고 과거 행동을 잊을 수 있습니다. Pure offline training은 넓은 history를 지키지만 publish가 느리고 delayed feedback을 다루기 어렵습니다. 독립 시스템 두 개를 운영하면 model dump, load, consistency boundary가 추가됩니다.
Primus mixed training runtime은 여러 batch와 stream source를 하나의 update process에서 결합합니다. Parameter update를 제어하고 source마다 세밀한 priority를 둡니다. Stream ingestion이 늦거나 load가 바뀌면 historical batch를 버리지 않으면서 fresh sample을 보호할 수 있습니다. Model은 offline-online serialization boundary를 별도로 넘지 않습니다.
운영 추천 모델 네 개의 offline evaluation AUC는 0.03~0.07% 개선됐습니다. Online advertising A/B test의 매출 개선은 0.4~2.4%였습니다. 작은 ranking 변화가 많은 auction에 영향을 줄 수 있고 revenue는 AUC에서 직접 변환되는 값이 아니므로 차이가 이상하지 않습니다. 이 범위를 infrastructure saving과 합쳐 평균을 내면 안 됩니다.
Freshness에는 bias와 feedback 위험도 있습니다. Source priority는 update마다 어느 population의 영향이 큰지 결정합니다. 늦게 들어온 conversion label은 과거 impression에 속할 수 있고 일시적인 traffic spike가 stream을 지배할 수 있습니다. Primus는 source를 섞는 메커니즘을 제공하지만 sampling, attribution, rollback 규칙은 business owner가 정해야 합니다.
가장 강한 근거인 5년 운영
Primus는 Douyin, Xigua, Toutiao를 포함한 ByteDance 추천 학습을 5년 동안 지원했습니다. 수천 모델이 플랫폼을 공유합니다. Scheduler와 storage가 바뀌는 동안 abstraction이 유지됐다는 점은 infrastructure layer에서 한 번의 benchmark보다 강한 근거입니다.
동시에 결과는 환경에 의존합니다. ByteDance는 scheduler, feature pipeline, resource telemetry, advertising feedback loop를 직접 운영합니다. 작은 조직은 horizontal pooling에 충분한 job이나 같은 control plane을 정당화할 source diversity가 없을 수 있습니다. 핵심 질문은 다른 배포 환경이 천만 core를 재현할 수 있는지가 아니라 공통 계약이 중복 integration을 줄이는지입니다.
Central architecture는 조직 전체 dependency가 될 수 있습니다. JobCRD나 DataCRD 변경은 여러 framework와 team에 영향을 줍니다. Compatibility, versioning, quota isolation, debugging tool이 planner speed만큼 중요합니다. 특이한 model이 private extension을 늘려 다시 fragment되지 않게 제한된 escape hatch를 제공해야 합니다.
서로 다른 세 분모의 구분
Primus는 infrastructure efficiency, control-plane speed, model outcome을 보고합니다. 17.1%는 core당 training throughput과 cluster cost, 23배는 task-graph generation time, 0.4~2.4%는 training mixture 변경 뒤 advertising revenue를 뜻합니다. 서로 다른 질문의 답이며 하나의 플랫폼 return으로 곱할 수 없습니다.
도입 검토에서는 먼저 기존 및 새 resource controller에 동일 job을 놓고 executor utilization과 completed sample을 측정해야 합니다. 그 다음 source 및 day count를 늘리며 planner latency를 시험합니다. Hybrid training은 freshness, model quality, business guardrail을 둔 별도 online experiment가 필요하며 마지막 단계만 product result를 귀속할 수 있습니다.
일반화할 교훈은 heterogeneous training이 resource, 데이터, update semantics의 독립적이면서 조합 가능한 contract로 관리된다는 점입니다. Primus는 데이터 logic을 다시 쓰지 않고 scheduler를 바꾸고, model code를 바꾸지 않고 store를 추가하며, 별도 training stack 없이 stream을 섞을 수 있습니다. 배치 규모는 역할 분리의 가치를 보여 주고 수치는 각 계층을 자체 분모로 평가해야 한다는 조건을 함께 제시합니다.
출처와 저작권 안내
이 글은 Silicon & Systems가 작성한 편집 분석으로, 구조와 측정, 한계를 우리 표현으로 다시 썼습니다. 원문의 문장, 표, 도판은 재수록하지 않았고 도판은 이 글을 위해 새로 만들었습니다. 전체 논문은 USENIX ATC 2025 발표 페이지에서 확인할 수 있으며, 저작권은 저자에게 있습니다. 2025.