Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models
Organizations: ECPI University
Abstract
Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs, and barriers to participation for researchers operating outside large industrial laboratories. This review examines the evolution of LLM scaling theory from empirical scaling laws to compute-optimal training, with particular emphasis on parameter efficiency, token utilization, data efficiency, and resource-constrained environments. Foundational work on scaling laws is synthesized alongside later research on compute-optimal training, data pruning, efficient architectures, quantization, low-rank adaptation, and edge-oriented optimization. The literature indicates a shift from scale maximization toward more deliberate allocation of parameters, tokens, compute, and hardware resources. At the same time, important empirical, theoretical, and methodological gaps remain regarding whether scaling principles established on enterprise-grade infrastructure generalize to smaller models and constrained computing environments. This review organizes these developments into a unified framework for resource-efficient LLM training and argues that future progress should evaluate efficiency not solely through model performance, but through the relationship among performance, parameter count, computational cost, token allocation, and hardware constraints.
Figures & tables
| Search term | Total results | Relevant studies | Date range |
|---|---|---|---|
| “LLM Training” AND (“Compute Overhead” OR “Resource Utilization”) | 371 | 90 | 2021–2025 |
| (“Parameter Efficiency” OR “Token Optimization”) AND “Scaling Laws” | 333 | 37 | 2021–2025 |
| “Compute-Optimal Scaling” AND (“Hoffmann et al.” OR “Scaling Efficiency”) | 77 | 1 | 2021–2025 |
| “Repeated Measures Design” AND (“AI Experiments” OR “Machine Learning”) | 1101 | 15 | 2021–2025 |
| “Edge Computing” AND (“Model Training” OR “Resource Constraints”) | 19872 | 15 | 2021–2025 |
| “Large Language Models” AND “Compute Efficiency” AND “Training Costs” | 44 | 19 | 2021–2025 |
| Study | Primary focus | Key insight | Efficiency relevance |
|---|---|---|---|
| Kaplan et al. (2020) | Empirical scaling laws | Loss scales predictably with model size, data, and compute. | Establishes the baseline relationship among resources and performance. |
| Bahri et al. (2021) | Theory of scaling laws | Statistical and geometric principles help explain observed power laws. | Adds theoretical structure to empirical scaling behavior. |
| Hutter (2021) | Learning-curve theory | Formal learning bounds provide an alternative perspective on performance growth. | Frames efficiency and learning limits theoretically. |
| Hoffmann et al. (2022) | Compute-optimal scaling | Model size and training tokens should be balanced under fixed compute. | Reorients scaling around token-to-parameter allocation. |
| Boopathy and Fiete (2024) | Scale-time equivalence | Smaller models trained longer can sometimes approach larger-model performance. | Challenges model-size dominance. |
| Sorscher et al. (2022) | Data-centric pruning | Data quality and selection can improve scaling efficiency beyond naive data growth. | Shifts attention from data quantity to data utility. |
| Model | Developer | Parameters | Estimated tokens |
|---|---|---|---|
| GPT-2 | OpenAI | 1.5B | 40B |
| GPT-3 | OpenAI | 175B | 300B |
| Gopher | DeepMind | 280B | 300B |
| MT-NLG | Microsoft/NVIDIA | 530B | 270–400B |
| Jurassic-1 | AI21 Labs | 178B | 300B |
| PaLM | 540B | 780B |
| Model | Developer | Parameters | Training tokens |
|---|---|---|---|
| Chinchilla | DeepMind | 70B | 1.4T |
| Phi-1 | Microsoft | 1.3B | 54B a |
| Phi-2 | Microsoft | 2.7B | 1.4T |
| Gemma 2B | 2B | 2T | |
| Gemma 7B | 7B | 6T | |
| TinyLlama | Open source | 1.1B | 1T |
| Dimension | Primary optimization question | Representative literature from the source review |
|---|---|---|
| Data/token | How much and what kind of data are needed to obtain useful learning? | Hoffmann et al. (2022) ; Sorscher et al. (2022) ; Pang et al. (2024) |
| Parameter | How much useful performance is obtained from model capacity and compute? | Maloney et al. (2022) ; Boopathy and Fiete (2024) ; Hoffmann et al. (2022) |
| Architecture | Can model structure reduce active compute or improve scaling behavior? | Tay et al. (2022) ; Wang et al. (2024) ; Sengupta et al. (2025) |
| Systems | Can memory, communication, synchronization, or numerical precision reduce realized cost? | Narayanan et al. (2021) ; Dettmers et al. (2023) |
| Hardware/edge | Can LLMs be trained, adapted, or served under explicit memory, power, or compute constraints? | Ahmed et al. (2025) ; Li et al. (2024) ; Yi et al. (2023) ; Yu et al. (2024) ; Moitra et al. (2025) ; Wei et al. (2025) |