GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement
Authors: Pavlos Stoikos, Foteini Oikonomou, Christos Poulos, Maria Pantazi-Kypriou, Athanasios Tziouvaras, Christos Anagnostopoulos, Georgios Karakonstantis, George Floros
Organizations: Department of Electrical and Computer Engineering, University of Thessaly, Volos, Greece · School of Computing Science, University of Glasgow, UK · Department of Electronic and Electrical Engineering, Trinity College Dublin, Ireland
Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particularly when combined with graph neural networks (GNNs) that capture circuit connectivity. However, most learning-based placement approaches focus on floorplanning, macro placement, or global placement, while detailed placement refinement remains relatively unexplored. In this paper, we present GPlaceRL, an open-source graph reinforcement learning framework for detailed placement refinement. GPlaceRL represents legalized placements as graphs and provides a modular environment for studying graph encoders, policy architectures, reward formulations, and local placement actions. To demonstrate the capabilities of GPlaceRL, we conduct a systematic evaluation of proximal policy optimization (PPO) policies with graph attention network (GAT) encoders in a per-design optimization setting. Across five placement benchmarks, the best greedy evaluation results achieve HPWL improvements ranging from 3.27% to 32.87%. The results highlight the importance of compact GAT architectures and flexible local action spaces for placement optimization. Overall, GPlaceRL provides a reproducible and extensible framework for systematic research on RL-based detailed placement refinement.
Figures & tables
Figure 1 : Overview of the GPlaceRL framework for detailed placement refinement. A legal placement is converted into a graph state and processed by a graph-based actor–critic policy. In the current implementation, a PPO-GAT policy applies local refinement actions, while periodic greedy evaluation selects the best checkpoint for generating the final optimized placement.
Design
Core size
Cells
Nets
IOs
Initial HPWL
B1
9.27×9.216
85
1494
26
1006
B2
21.78×21.888
437
12128
55
13028
B3
54.0×54.144
2976
135314
632
206641
B4
76.14×76.032
6848
372952
332
723181
B5
129.24×129.024
19463
1010888
674
3469054
Table I : Statistics of the evaluated placement benchmarks.
Design
Action space
Best eval Δ (%)
Best config
Best episode
B1
Move only
28.26
L2-H4
1400
B1
Swap only
23.10
L2-H4
60
B1
Swap + Move
32.87
L2-H4
1360
B2
Move only
12.65
L2-H8
1380
B2
Swap only
8.86
L4-H8
70
B2
Swap + Move
13.35
L2-H8
1280
Table II : The top-performing RL architecture for each benchmark and action space.
Figure 2 : Qualitative placement and wire-density analysis for the best B3 result in Table II , obtained with the swap & move action space and yielding a 12.11% HPWL improvement. (a) Initial placement with selected local windows. (b) Refined placement after optimization. (c) Before-optimization wire-density maps for the selected windows. (d) After-optimization wire-density maps. (e) Density-change maps computed as before minus after, where red indicates reduced wire density and blue indicates increased wire density.
Figure 3 : Convergence behavior for the B4 swap & move experiment. The best evaluation reaches a 5.25% HPWL improvement at episode 420 , corresponding to the selected B4 result in Table II .
Chip placement is a critical step in physical design. While reinforcement learning (RL)-based methods have recently emerged, their training primarily focuses on wirelength optimization, and therefore often fail to achieve expert-quality layouts. We identify the reward design as the primary cause for the performance gap with experts, and instead of formalizing intricate processes, we circumvent this by directly learning from expert layouts to derive a reward model. Our approach starts from the final expert layouts to infer step-by-step expert trajectories. Using these trajectories as demonstrations or preferences, we train a model that captures the latent implicit rewards in expert results. Experiments show that our framework can efficiently learn from even a single design and generalize well to unseen cases.
Ruo-Tong Chen, Ke Xue, Chengrui Gao +7
State Key Laboratory of Novel Software Technology, Nanjing University, China · School of Artificial Intelligence, Nanjing University, China · Huawei Noah’s Ark Lab, China
The application of Reinforcement Learning (RL) in Electronic Design Automation (EDA), particularly for chip placement, has attracted considerable attention in recent years. While existing machine learning (ML)-based approaches have achieved notable progress, they predominantly focus on generating optimal layouts in a single attempt, often producing solutions that require subsequent refinement. To address this limitation, we propose RollPlace, a novel and generalized macro placement framework. RollPlace adopts a two-stage optimization strategy: generating initial placement solutions via machine learning methods or heuristic-based strategies, and refining these layouts efficiently by adjusting specific macros derived from the initial stage. This strategy circumvents the sequential generation constraints inherent in traditional RL-based placement methods. Furthermore, RollPlace seamlessly integrates Monte Carlo Tree Search (MCTS) to balance exploration and exploitation, and employs a rollout mechanism for efficient local search. Extensive experiments on the ISPD 2005 benchmark demonstrate that RollPlace outperforms state-of-the-art methods. Additionally, end-to-end experimental results based on OpenROAD across 19 benchmarks show that RollPlace excels in multiple metrics. The proposed framework offers a robust and scalable solution for addressing the growing complexity of modern chip design challenges.
Qi Zhou, Guojun Liu, Guangzhi Qi +5
Faculty of Computing, Harbin Institute of Technology, Harbin, 150001, China · School of Materials Science and Engineering, Harbin Institute of Technology, Harbin, 150001, China
Routing and switch placement are fundamental combinatorial optimization problems in chip design, requiring the joint optimization of routing topology and physical placement under strict structural, geometric and logical constraints. Existing approaches typically rely on carefully engineered heuristics that incorporate strong problem-specific biases to navigate the enormous space of possible designs. In this work, we introduce a hierarchical reinforcement learning framework for joint routing and switch placement at the level of logical communication routes. Starting from a minimal routing graph, our method progressively constructs increasingly expressive solutions through three coupled operations: switch expansion, switch placement, and route refinement. These operations preserve routing validity by construction, restricting exploration to feasible configurations where every communicating initiator-target pair has one assigned loop-free route. We explore the induced solution space using Gumbel Monte Carlo Tree Search, showing that neural-guided search substantially improves solution quality over non-learning optimization methods. Furthermore, pretraining across floorplans provides a strong initialization for fine-tuning on unseen instances.
Dorian Gailhard, Ugo Lecerf, Enzo Tartaglione +2
LTCI, Télécom Paris, Institut Polytechnique de Paris, France · Arteris IP