The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes THEIA, a novel dataset containing thousands of layout images paired with question-answer conversations, along with a benchmark that employs a fine-tuned vision-language model (VLM) to analyze GDSII files of analog circuits, enabling designers to interact with and query physical layouts as intuitive, meaningful entities. Experimental results using thousands of analog designs across five realistic tasks demonstrate that the proposed fine-tuned VLM outperforms state-of-the-art general-purpose VLMs by a significant margin (up to 73%), highlighting a fundamental gap between general-purpose multimodal reasoning and domain-specific layout understanding.
Figures & tables
Figure 1
Figure 3 : Methodology overview. It comprises three main phases: (i) Dataset Creation , (ii) VLM Fine-tuning , and (iii) VLM Deployment . Dataset Creation generates sets of task-specific questions and answers ( Q&As ) from a collection of GDSII layouts ( GDS ). VLM Fine-tuning uses training and validation splits to fine-tune a VLM to the target domain. VLM Deployment evaluates the fine-tuned model on the held-out test data, employing an LLM-as-a-judge to assess prediction correctness.
Figure 4 : Detailed methodology stages: (i) Dataset Creation (top), (ii) VLM Fine-tuning (middle) and (iii) VLM Deployment (bottom).
Model
Family
Params
LB
VE
MM Train.
Inst-tuned
Res.
Sp. Aw.
GPT-5.2 OpenAI (2025)
OpenAI
–
Proprietary
Proprietary
Gemini-3-Flash Google (2025)
Google
–
Proprietary
Proprietary
GLM-4V-9B GLM et al. (2024)
Zhipu AI
9 B
GLM
ViT
InternLM-4KHD Dong et al. (2024)
InternLM
7 B
InternLM
ViT (4K)
InternVL2 Chen et al. (2024b)
InternVL
8 B
InternLM
Hybrid
InternVL3.5 Wang et al. (2025)
InternVL
8 B
InternLM
Hybrid
Table 1 : VLMs considered in this work, compared by backbone, vision encoder, multimodal and instruction tuning, supported resolution, and spatial awareness. InternLM-4KHD and Qwen2.5-VL-7B are fine-tuned for the controlled backbone-transfer study.
Category
Circuit Type
Variants (#)
Avg. Devices (#)
NMOS (#)
PMOS (#)
CAP (#)
RES (#)
Total Devices (#)
Single Device
Capacitor
4997
1
–
–
1
–
4997
NMOS Transistor
5000
1
1
–
–
–
5000
PMOS Transistor
5000
1
–
1
–
–
5000
Resistor
5000
1
–
–
–
1
5000
Subtotal
19997
–
–
–
–
–
19997
Base Circuits
Ahuja OTA
995
15
10
4
1
–
14925
Table 2 : Analog circuits dataset composition and statistics. For each category, the table reports the circuit types, the number of layout variants, the average number of devices per variant, and the average per-device-type counts (NMOS, PMOS, capacitors, and resistors).
Complexity
Task ID
Task Description
Q&As (#)
Easy
A
Identification of single component devices (capacitors, resistors, NMOS, PMOS)
19997
Medium
B
Identification of base circuit topologies (OTAs, filters, regulators, gate drivers)
5894
C
Component counting and enumeration in base circuits
27475
Hard
D
Component counting and enumeration in complex mixed circuits
19848
E
Identification of base circuit topologies within complex mixed circuits
4140
Total
77354
Table 3 : Task descriptions, organized by complexity (i) Easy, (ii) Medium, and (iii) Hard.
Figure 5 : Examples of GDSII view of single-component devices layouts.
Complexity
Task
Count (#)
InternLM FT
InternLM no FT
GPT-5.2
Qwen2.5-VL-7B
Pass@1
Pass@5
no FT
FT
Easy
A
1001
100%
100%
41%
100%
19%
100%
Medium
B
300
84%
92%
16%
83%
17%
85%
C
1399
83%
94%
18%
20%
23%
93%
Hard
D
1003
63%
87%
24%
26%
24%
89%
E
207
81%
89%
7%
10%
20%
91%
Table 4 : Accuracy on the held-out THEIA test splits. The Qwen2.5-VL-7B transfer uses the same splits, prompts, answer format, and deterministic decoding with and without task-specific fine-tuning.
Task
THEIA test
ALIGN test
Δ
B: topology identification
84.0% ( 252/300 )
89.5% ( 34/38 )
+5.5 pp
C: component counting
83.0% ( 1161/1399 )
73.7% ( 84/114 )
−9.3 pp
Table 5 : Cross-generator evaluation of InternLM-4KHD fine-tuned on THEIA and evaluated on held-out ALIGN Kunal et al. (2019) layouts.
Task
Judge Pass@1
EM
NSM
Count Acc.
Topology-set F1
A
100%
100%
100%
–
–
B
84%
84%
84%
–
–
C
83%
83%
83%
83%
–
D
63%
63%
63%
63%
–
E
81%
81%
81%
–
89%
Table 6 : Deterministic validation metrics for fine-tuned InternLM-4KHD on held-out THEIA samples. LLM-judge Pass@1 is shown for comparison.
Task
Count
Cohen’s κ
Accuracy
Precision
Recall / F1
A
1001
1.0
100%
100%
100% / 100%
B
300
1.0
100%
100%
100% / 100%
C
1399
1.0
100%
100%
100% / 100%
D
1003
1.0
100%
100%
100% / 100%
E
207
1.0
100%
100%
100% / 100%
Table 7 : Agreement of Qwen3-32B LLM judge with deterministic evaluation on held-out samples.
Figure 6 : Accuracy trends as the amount of task-specific fine-tuning data is progressively increased.
Task
Multiple choice
Open ended
A
100%
100%
B
84%
83%
C
83%
99%
D
63%
75%
Table 8 : Fine-tuned InternLM-4KHD accuracy for multiple-choice and open-ended forms of Tasks A-D on held-out layouts.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value / Default
Meaning / Notes
A. Base model and task formulation
base_model
internlm/internlm- xcomposer2-4khd-7b
Pretrained VLM used as initialization (loaded with custom model code).
task_type
causal language modeling
Supervised instruction tuning in a chat-style transcript format.
data_format
single-turn, single-image
Each sample contains one image and one question-answer pair (serialized with role and boundary tokens).
B. Optimization and training schedule
num_train_epochs
1
Number of training epochs.
Appendix
Table 9 : Fine-tuning and inference parameters.
Figure 7 : Dataset structure.
Complexity
Task
Train (#)
Validation (#)
Test (#)
Total (#)
Circuit Type Distribution
Easy
A
17997
999
1001
19997
NMOS: 5000 ( 25.0% )
PMOS: 5000 ( 25.0% )
RES: 5000 ( 25.0% )
CAP: 499 7 ( 25.0% )
Medium
B
5302
292
300
5894
Gate Driver: 1000 ( 17.0% )
Ahuja OTA: 995 ( 16.9% )
Appendix
Table 10 : Subsets splits cardinality and balance analysis.
Layout area ( μm2 )
Polygon count (#)
Layers (#)
Aspect ratio (-)
Routing density (-)
Topology
p25
med
p75
p25
med
p75
p25
med
p75
p25
med
p75
p25
med
p75
Dataset groups
Single Devices
94
125
663
43
51
228
8
9
10
0.6
0.6
2.1
5.5
5.9
6.5
Base Circuits
2965
6149
9828
4073
4674
5471
13
14
14
1.0
1.5
2.1
3.7
4.0
4.6
Mixed
6209
10545
14887
9407
13109
17365
14
14
15
0.7
1.3
2.0
3.5
3.8
4.3
Single Devices
Appendix
Table 11 : Dataset statistics across groups and circuit topologies.
Task
Layout
Question and choices
Target
A
single_cap_131
What type of device is this? A. NMOS; B. PMOS; C. resistor; D. capacitor. Pick one answer only.
D (capacitor)
B
ahuja_ota_481
What circuit is shown? A. gate driver; B. LDO; C. Miller OTA; D. HPF; E. Ahuja OTA; F. LPF. Select exactly one option.
E (Ahuja OTA)
C
ahuja_ota_308
Can you count the transistors in this circuit? A. 12; B. 14; C. 11; D. 15. Select exactly one option.
B (14)
D
mixed_8977
Can you count the PMOS devices? A. 10; B. 9; C. 8. Select exactly one option.
C (8)
E
mixed_661
What base circuits are combined? A–F enumerate topology multisets; the correct option is F: one HPF, one LDO, and one Miller OTA. Select exactly one option.
F
Appendix
Table 12 : Representative held-out Q&As retrieved from the released THEIA task files. The answer letter is the assistant target in the corresponding JSON record.
Task
3 options
4 options
5 options
6 options
C ( 1399 )
370
359
349
321
D ( 1003 )
272
258
223
250
Appendix
Table 13 : Number of answer options for Task-C/D held-out questions after distractor construction.
Topology
Train ( 5302 )
Validation ( 292 )
Test ( 300 )
Ahuja OTA
895 ( 16.9% )
49 ( 16.8% )
51 ( 17.0% )
Gate Driver
900 ( 17.0% )
50 ( 17.1% )
50 ( 16.7% )
HPF
865 ( 16.3% )
48 ( 16.4% )
49 ( 16.3% )
LDO
890 ( 16.8% )
49 ( 16.8% )
50 ( 16.7% )
LPF
873 ( 16.5% )
48 ( 16.4% )
50 ( 16.7% )
Miller OTA
879 ( 16.6% )
48 ( 16.4% )
50 ( 16.7% )
Appendix
Table 14 : Task-B class counts (percentage within split) and majority-class baseline.
Figure 8 : Example of a base circuit layout, an HPF. Devices are highlighted in coloured boxes.
School of Integrated Circuits, Peking University, Beijing, China · Beijing Advanced Innovation Center for Integrated Circuits, Beijing, China · Institute of Electronic Design Automation, Peking University, Wuxi, China