Organizations: Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong · Huazhong Agricultural University, Wuhan, China
Proactive power management systems reduce processor dynamic power through runtime power prediction and power-aware scheduling. Accurate, stable and low-overhead digital on-chip power meters (OPMs) are crucial for improving the prediction quality. Recent studies have explored various modeling methods, including using linear models, decision trees, and multi-layer perceptrons (MLPs) to construct OPMs. However, most current approaches train models end-to-end without analyzing the physical interpretability of features, affecting their ability to generalize to unseen workloads. Grounded in the design principles of synchronous digital VLSI circuits, X-OPM introduces a robust feature engineering framework that uses tree-based models to capture feature interactions and linear models for prediction. It also incorporates a human-in-the-loop workflow to balance model accuracy against modeling effort. Evaluated on a commercial C906 vector processor, X-OPM consistently achieves R2>0.93 across all workloads with sampling window size set below 8 cycles. In contrast, state-of-the-art methods including APOLLO, COBIT, and standard MLPs fail to generalize across all test cases. Layout with commercial EDA tools shows that X-OPM incurs an area overhead below 0.1%, which is on par with lightweight tree-based and linear models, and significantly smaller than MLP-based models.
Figures & tables
Fig. 1: Board TDP (Thermal Design Power) versus release year of recent cloud AI processors.
Fig. 2: A typical single-stage synchronous circuit design for analyzing the source of dynamic power.
Category
Naming Convention
Post-processing *
Power Behavior
Control
*_busy *_stall *_clk_en
Split signals into single-bit and invert the one that is negatively correlated with power
Positively correlated with power
Config.
*_mode *_sel *_cfg
Decimal to one-hot
Non-linear and requires feature interaction
Data
*_data *_bus
Calculates Hamming distance ( HD ) between consecutive cycles
HD is positively correlated with power
* Binning can be applied to all categories.
TABLE I: The proposed three-category signals.
Fig. 3: Dynamic power fluctuation of a two-stage pipeline. (A) Block Diagram of the pipeline (B) Data toggle at the output of two registers. (C) Per-cycle dynamic power decomposition.
Category
Verification Purpose
Example
Direct
Stimulate a target module or function to check its correctness
Issue one instruction to check the functional correctness of the processor
Random
Randomize simulation parameters to uncover corner-case bugs
Randomly issue instructions to check the in-order retirement of an out-of-order processor
Benchmark
Check the architectural model’s performance matches with the RTL implementation
CoreMark [ 25 ] for CPUs or a complete neural network inference for GPUs
TABLE II: Categories of pre-silicon test cases.
Fig. 4: The proposed automatic modeling pipeline
Fig. 5: Architecture of the OPM predictor and its model evaluation datapath.
Fig. 6: Power breakdown of the C906 core running the CoreMark benchmark. Bar heights show average module power; error bars show the power fluctuation range during the execution.
Fig. 7: The residual plot shows the predicted versus true per-cycle power of x_aq_core on the test benchmarks. The four baselines on the left are trained on raw single-bit features; X-OPM is trained on the processed features. Each hexagon bin is colored by the number of cycles it holds; the red dashed line marks perfect prediction. Train and validation R2 are listed in the lower-right corner of each panel.
Fig. 8: Test R2 of the C906 OPM versus the number of retained proxies under RFE pruning.
Accurate power estimation is important for understanding and optimizing CPU power behavior, yet practical workflows often rely on simulation-derived information or post-silicon analysis. In this work, we present BigPower, a hierarchical source-level surrogate model for fine-grained module-level power estimation during CPU design. BigPower leverages large language model-based representations together with architectural hierarchy, module connectivity, configuration parameters, and workload context to estimate module-level power consumption directly from source-level design information, without requiring additional simulation during inference. Experimental results in the open-source XiangShan processor family demonstrate practical fine-grained power estimation across diverse configurations and workloads, offering an efficient alternative to conventional simulation-based workflows.
Honghua Zhu, Chunjie Luo, Jianfeng Zhan
State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important for both economic and environmental reasons. Traditional methods for power modeling are inadequate in these dynamic software-defined environments due to their inability to model complex and nonlinear factors affecting energy use. We investigate the use of feature extraction and regressor-based machine learning methods for predicting power consumption in virtualized open radio access networks (O-RANs), utilizing datasets from a hardware-instrumented testbed. We test three variants of deep neural networks (DNNs), namely, a standard DNN, a regularized DNN, and a hybrid model combining DNN-based feature extraction with an XGBoost regressor. We evaluate the performance of these models for various system parameters such as transmission gain, modulation/coding schemes, and airtime. We show that the hybrid model consistently outperformed others, achieving a mean relative error below 0.5%. Results suggest hybrid models like DNN-XGBoost offer superior accuracy and could be integrated into O-RAN management tools to enable more energy-efficient network orchestration in future networks.
Estimating CPU power on heterogeneous ARM-based commodity devices is challenging due to limited access to CPU's voltage domains. As a result, state-of-the-art energy-aware Federated Learning (FL) frameworks typically rely on simplified approximate power models to estimate computation energy, rather than the more accurate analytical CMOS-based model. To bridge this gap, we propose a reproducible CPU power estimation methodology combined with a rail-to-cluster mapping technique to retrieve cluster-level supply voltage. We evaluate our approach on two commodity Android devices and show that the analytical model predicts CPU power with errors below 10%, whereas the approximate model incurs errors of up to 959%. Using AnycostFL, a state-of-the-art energy-aware FL framework, we show that the analytical model achieves the same 80% model accuracy while consuming 1.4x less energy than the approximate model. These results highlight that approximate models can severely misestimate computation energy and lead to suboptimal decisions. This work facilitates the use of analytical CPU power models on heterogeneous multi-cluster ARM-based mobile SoCs without additional hardware support or external power measurement tools.
Chaimae Jallouli, Karim Boubouh, Robert Basmadjian
Mohammed VI Polytechnic University, Benguerir, Morocco · Khalifa University, Abu Dhabi, UAE