Training Compute-Optimal Transformer Models: A Review of Scaling Laws, Allocation, and Practice

Alexander Memming

SSRN Electronic Journal · 2026

The performance of transformer models improves predictably with the scale of three coupled resources: the number of model parameters, the volume of training data, and the total compute budget consumed during training. A central practical question is how to allocate a fixed compute budget between model size and data size so as to minimize loss-the compute-optimal training problem. This review synthesizes the empirical and theoretical literature on neural scaling laws and compute-optimal training of transformers.

We trace the development from early evidence that generalization error follows power laws in scale, through the influential parameter-centric prescription of Kaplan et al. [2020], to the data-balanced reallocation of Hoffmann et al. [2022] (Chinchilla) and its subsequent replications and refinements. We then survey the directions that have reshaped the compute-optimal frontier in practice: accounting for inference cost, training under data constraints and with repeated data, sparse mixture-of-experts architectures, the (limited) role of architectural inductive bias, and the generalization of scaling laws across modalities and scientific domains. We close with a discussion of methodological pitfalls, open problems, and the practical recipe that the field has converged on.

Throughout, we emphasize that scaling laws are an empirical instrument for budget allocation and risk reduction rather than a law of nature, and that their coefficients are contingent on data, tokenizer, optimization, and evaluation choices.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.