AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance
Authors: Junyu Deng, Jiale Cao, Mengtian Li, Zhongxia Ji, Ruhua Chen, Yiyi He, Guangnan Ye, Zuo Hu
Organizations: Fudan University, Shanghai, China · Shanghai University, Shanghai, China · Shanghai Film Academy and Shanghai University, Shanghai, China · Shanghai Theatre Academy, Shanghai, China · Nanjing University of the Arts, Nanjing, China
We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling. Lighting design in live performance settings requires the seamless translation of musical features into dynamic lighting behaviors. However, traditional workflows remain time-consuming, labor-intensive, and difficult to transfer. AuraLuxMuse encodes music and professional cue sequences into a shared retrieval space, estimates cue-event density, and retargets selected fixture commands to the destination stage. It assists pre-production authoring by returning editable cues rather than replacing the designer with an unconstrained generator. At the heart of AuraLuxMuse are two key modules: Lighting-Aligned Music Pretraining (LAMP), which performs contrastive learning between audio and lighting cues for alignment, and Preference-Adaptive Mixture of Experts (PAMoE), which conditions preference-aware cue retrieval and adaptation on designers' intent through a gated ensemble of style-specific expert networks. To support training and evaluation, we introduce Musilux, the first dataset of paired musical audio and professional lighting cue sequences under diverse performance scenarios. We evaluate AuraLuxMuse across both virtual simulation environments and professional-grade laboratories. Experimental results, including objective and subjective evaluation, demonstrate that AuraLuxMuse retrieves and adapts stage-lighting cues that are visually cohesive, semantically meaningful, and artistically expressive, showing its potential for AI-assisted aesthetic stage design.
Figures & tables
Figure 1. Procedure Comparison: Traditional Stage Lighting Design vs. AuraLuxMuse. Conventional pipelines require designers to manually decompose music into artistic components, then iteratively design cues, align timing, and input parameters into consoles. AuraLuxMuse accelerates this workflow through feature extraction, preference-conditioned retrieval, cue adaptation, and console dispatch. Comparison of a traditional manual stage-lighting workflow with the AuraLuxMuse workflow, from music input and artistic analysis to console-ready lighting output.
Figure 2. Overview of the AuraLuxMuse Pipeline. Music and lighting data are individually preprocessed before entering LAMP for contrastive learning, and preference representations are injected via PAMoE. P1/P2/P3 denote per-style frame counts, and Ti=∑t=1i∣Pt∣ is the cumulative sample count used to index the illustrated segments. Cosine similarity and the metadata value K predicted by the non-causal TCN Metadata Predictor are used to query the cue corpus. AuraLuxMuse pipeline showing music and cue encoders, LAMP contrastive alignment, PAMoE preference routing, the non-causal TCN Metadata Predictor, top-K cue retrieval, and console execution.
Figure 3. Musilux Construction Procedure. Music and lighting cues are collected, aligned both within cue structures and across modalities, and then aggregated into the Cue Sequence Corpus that constitutes the dataset. Flow diagram of the Musilux dataset construction procedure: collection, intra-cue alignment, cross-modal alignment, and aggregation into the Cue Sequence Corpus.
Figure 4. Qualitative Results. AuraLuxMuse and the manual design exhibit different spatial and temporal choices in the same Drama scenario; these examples illustrate scenario-dependent competitiveness rather than universal superiority. Side-by-side qualitative comparison frames between manually designed lighting and AuraLuxMuse-generated lighting in a Drama scenario, illustrating temporal continuity.
Figure 5. Stage Demonstrations across Lighting Styles. Lighting configurations are visualized across varying levels of Color harmony, Brightness control, and Zone distribution. Definitions of stage-lighting terms are provided in parentheses. Grid of stage demonstrations showing different combinations of color harmony, brightness levels, and zone distributions for the lighting style space.
Scenario
Source
FID
Sexp↑
Snon
EGSS ↑
Concert
Manual
8.21
2.58 (1.33)
4.31 (0.73)
3.09 (1.15)
AuraLuxMuse
3.65 (1.00)
3.69 (0.88)
3.66 (0.75)
Drama
Manual
4.47
3.08 (1.10)
4.03 (0.93)
3.36 (1.04)
AuraLuxMuse
2.71 (1.24)
3.27 (1.11)
2.88 (0.93)
Entertainment
Manual
2.00
2.46 (1.25)
3.31 (0.98)
2.71 (1.16)
AuraLuxMuse
3.21 (1.25)
3.39 (1.13)
3.26 (0.94)
Table 1. Quantitative Results. Average expert and non-expert ratings are reported for all scenarios. AuraLuxMuse achieves scenario-dependent competitiveness with the manual counterparts. Values in parentheses are standard deviations. FID measures the representational discrepancy between AuraLuxMuse and the corresponding reference.
Strategy
FID ↓
Precision@5 ↑
mAP ↑
EGSS ↑
Random
6.32
0.07
0.03
2.71 (0.72)
Position-based
25.14
0.07
0.04
1.98 (0.73)
Frequency-based
16.32
0.07
0.03
2.44 (0.76)
Nearest-neighbor
21.65
0.08
0.04
2.10 (0.73)
Rule-based
6.61
0.06
0.02
2.81 (0.81)
Similarity-based
4.89
0.17
0.10
3.27 (0.88)
Table 2. Comparison of Retrieval Strategies. We compare similarity retrieval with random, position-based, frequency-based, nearest-neighbor, and rule-based strategies. Values in parentheses are standard deviations.
Table 3. Subjective Results across Music Styles. EGSS is reported for six validation styles and six test styles. Validation uses held-out examples with manual references; test tracks and the test stage are excluded from training.
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1 . Qualitative Results for AuraLuxMuse Variants in the Drama Scenario. Three architecture variants produce distinct retrieved/adapted lighting outputs as editable references for professional designers. A grid of stage-lighting frames comparing manual design with three AuraLuxMuse variants in the Drama scenario.
Figure S2 . Qualitative Results in the Concert Scenario. Consecutive frames compare manual lighting with AuraLuxMuse retrieved/adapted outputs under the same scenario; the variants show how the model structure changes cue selection. A grid of stage-lighting frames comparing manual design with AuraLuxMuse variants in the Concert scenario, together with music features and event timestamps.
Figure S3 . Qualitative Results in the Entertainment Scenario. Consecutive frames compare a manual design with AuraLuxMuse retrieved/adapted cues under the same Entertainment scenario. The example illustrates different choices in visual complexity and atmosphere without implying universal superiority. A grid of consecutive Entertainment-scenario lighting frames comparing manual design with AuraLuxMuse retrieved and adapted cues, plus music features and event timestamps.
Ratings
Color Harmony
1
Harsh, unbalanced colors; poor harmony
2
Colors clash; mildly unbalanced
3
Colors blend acceptably; moderate harmony
4
Colors blend smoothly; strong harmony
5
Colors blend flawlessly; outstanding harmony
Appendix
Table S1 . Lighting Evaluation for Experts: Color Harmony. This dimension rates the aesthetic balance and blending of stage-lighting colors across the complete visual composition of each scene.
Ratings
Brightness Control
1
Uneven brightness; excessively dim/bright
2
Uneven brightness; slightly harsh
3
Reasonable brightness control; minor issues
4
Consistent brightness; well-adjusted
5
Perfect brightness; optimal balance
Appendix
Table S2 . Lighting Evaluation for Experts: Brightness Control. This dimension rates the consistency and adjustment of brightness across an entire performance sequence under review.
Ratings
Zone Distribution
1
Uneven zones; inadequate coverage
2
Uneven zones; noticeable gaps
3
Fair zoning; slight inconsistencies
4
Even zoning; effective coverage
5
Flawless zoning; complete coverage
Appendix
Table S3 . Lighting Evaluation for Experts: Zone Distribution. This dimension rates evenness and coverage across the visible stage area.
Ratings
Rhythmic Tempo
1
No rhythm; disorganized timing
2
Weak rhythm; moderately off
3
Average rhythm; reasonably aligned
4
Good rhythm; accurately timed
5
Excellent rhythm; impeccably timed
Appendix
Table S4 . Lighting Evaluation for Experts: Rhythmic Tempo. This dimension rates lighting synchronization with performance rhythm.
Ratings
Narrativity
1
Fails to convey plot / emotion
2
Weakly conveys plot / emotion
3
Moderately conveys plot / emotion
4
Clearly conveys plot / emotion
5
Strongly conveys plot / emotion
Appendix
Table S5 . Lighting Evaluation for Experts: Narrativity. This dimension rates how effectively lighting conveys plot and emotion.
Ratings
Evaluation Dimensions
Rhythmic Alignment
Thematic Suitability
Visual Consistency
1
No alignment
Not suited
Extremely inconsistent
2
Poor alignment
Poorly suited
Somewhat inconsistent
3
Basic alignment
Basically suited
Basically consistent
4
Fair alignment
Fairly suited
Generally consistent
5
Perfect alignment
Perfectly suited
Always consistent
Appendix
Table S6 . Lighting Evaluation Dimensions (Relative to Music) for Non-Experts. The criteria capture spectators’ intuitive perceptions of performance quality across the three evaluated scenarios.
Peripheral, Atmospheric, Neutral, Attention, Dominant Zone
Rhythmic Tempo
Static, Slow, Moderate, Fast, Rapid
Appendix
Table S7 . Overview of Lighting Style Dimensions. Color, Brightness, Zone, and Rhythm are mapped to representative attributes.
Level
Description
Illuminance (lux)
Dark
Almost invisible
0–20
Dim
Barely visible
20–100
Moderate
Visually comfortable
100–300
Bright
Clear visibility
300–700
Intense
Extremely bright
700+
Appendix
Table S8 . Brightness Attribute Levels. Each level is defined by a perceptual description and illuminance range for preference specification.
Level
Description
Peripheral
Invisible boundary or background layer
Atmospheric
Subtle lighting for mood creation
Neutral
Balanced fill without drawing attention
Attention
Guides audience attention
Dominant
Strongest visual focus
Appendix
Table S9 . Zone Distribution Attribute Levels. Each level defines a degree of visual prominence within a scene.
Level
Description
Static
Minimal change, nearly static
Slow
Gentle & Sparse transitions, rare rhythmic cues
Moderate
Regular tempo, structured but not fast
Fast
Frequent changes, strong rhythmic presence
Rapid
Intense flashes with strong beats
Appendix
Table S10 . Rhythmic Tempo Attribute Levels. Each level defines a transition pace for preference specification.
Figure S4 . Patching Mechanism Overview. Retrieved artistic cues are retargeted to compatible fixtures through a logical layout map and an industrial-console patch table used for dispatch to the console. A schematic showing retrieved lighting cues being retargeted through compatibility checks, a logical layout map, and a patch table before console control.
Music Encoder / Cue Encoder
Transformer
ResNet
VGGish + MLP
0.699( ↓ )
2.015
VGGish + ResNet
2.181
3.926
Appendix
Table S11 . Ablation of LAMP Architectures. Contrastive loss is reported after 500 training epochs of model optimization.
Model Variant
FID (↓)
Sexp
Snon
EGSS (↑)
PAMoE Only
18.06
2.26
2.86
2.44
Soft Only
10.97
2.83
3.47
3.02
PAMoE + Soft
4.89
3.19
3.45
3.27
Appendix
Table S12 . Ablation of PAMoE and Soft Alignment. All experiments use the same learning rate, weight decay, and number of epochs.
Architecture
MAE ↓
RMSE ↓
R 2 ↑
Pearson r↑
Bias ↓
MLP
16.987
24.113
-368.715
0.061
+16.973
TCN Density
0.343
0.449
0.872
0.934
+0.017
Appendix
Table S13 . Metadata Predictor Performance. We compare window-level cue counting with a baseline MLP and our non-causal TCN.
Architecture
FID ↓
EGSS ↑
Precision@5 ↑
mAP ↑
CLAP
5.06
2.52
0.11
0.043
SLAP
7.20
2.11
0.02
0.085
AudioCLIP
5.21
2.18
0.13
0.05
AuraLuxMuse
4.89
3.27
0.170
0.100
Appendix
Table S14 . Comparison with Contrastive Baselines. AuraLuxMuse obtains the highest Precision@5 and EGSS under the reported setup.
Architecture
FID
EGSS ↑
Inference Time (s/file) ↓
T5
4.28
1.79
25.74
BART
4.21
1.89
35.20
Autoformer
3.98
1.24
207.27
SpecTNT
3.95
1.24
150.18
AuraLuxMuse
4.89
3.27
7.65
Appendix
Table S15 . Comparison with Generative Baselines. Against four sequence-to-sequence models, AuraLuxMuse achieves lower inference latency and a higher EGSS under the reported setup.
Strategy
K
FID ( ↓ )
EGSS ( ↑ )
Fixed Small
1
8.42
2.85
Fixed Large
10
9.82
3.02
Adaptive
Dynamic
4.89
3.27
Appendix
Table S16 . Ablation Study: Static vs. Dynamic K Values. Dynamic K yields a lower FID and a higher EGSS than either static setting.
Soft Weight ( w )
Avg FID ↓
R@1 ↑
R@5 ↑
R@10 ↑
MRR ↑
nDCG ↑
w=0.1
11.85
0.045
0.223
0.338
0.144
0.303
w=0.3
4.89
0.046
0.249
0.374
0.149
0.309
w=0.5
10.10
0.062
0.202
0.280
0.143
0.300
w=0.7
61.21
0.033
0.159
0.301
0.118
0.279
w=0.9
47.09
0.020
0.127
0.258
0.097
0.261
Appendix
Table S17 . Ablation on Soft Alignment Weight ( w ). We evaluate the impact of the soft-alignment weight w on cross-modal retrieval performance and Fréchet Inception Distance (FID). The hard-alignment weight is fixed at 1.0. Results indicate that w=0.3 yields the best balance, achieving the strongest overall retrieval metrics by leveraging stylistic similarities without blurring discriminative boundaries.
Embedding
FID (↓)
Sexp
Snon
EGSS (↑)
One Hot
12.24
2.51
3.29
2.74
Attribute Map
4.89
3.19
3.45
3.27
Appendix
Table S18 . Ablation of Cue-Embedding Strategies. The architecture is fixed, and only the embedding strategy varies.
Figure S5 . Sensitivity analysis of the EGSS metric with respect to λ . The parameter λ controls the weight assigned to expert preferences in the subjective evaluation process. The manual lighting design serves as an approximate upper bound. AuraLuxMuse consistently outperforms all automated baselines across the full λ range, demonstrating strong robustness to this hyperparameter. The vertical dashed line indicates our selected default configuration ( λ=0.7 ). A line chart showing EGSS for manual design and automated baselines as the expert-preference weight lambda varies, with a dashed line marking lambda equal to 0.7.
Figure S6 . Examples of Stage Lighting Fixtures: (a-b) Imaging Lights: FINE 300C LEKO and FINE 400T/D LEKO for sharp projection and patterned illumination; (c-d) Spotlights: FINE 420 BEAM IP and FINE 480 BSW IP for dynamic focus and concentrated beams; (e-f) Soft Lights: FINE 600T/D PANEL and FINE 300T/D PANEL for diffused illumination and even coverage; (g-h) Floodlights: FINE 1514 DG and FINE 1514 ZOOM for broad lighting and wide effects. An eight-panel image sheet showing imaging lights, spotlights, soft lights, and floodlights used in the stage-lighting setup.
Figure S7 . Fixture Parameter Visualization. The Fixture Sheet records instantaneous lighting-infrastructure states. Its fields match the functional components in Section S4 , enabling precise state tracking during acquisition. A fixture-sheet table listing the instantaneous state parameters of the lighting infrastructure for data acquisition and tracking.
Figure S8 . Software-Defined Spatial Manifold: Top-down View. The top-down orthographic projection of the experimental environment as registered within the GrandMA2 configuration space. Each luminaire instance and structural element is precisely localized within a unified global coordinate system, providing the necessary extrinsic parameters for spatiotemporal alignment. This geometric blueprint serves as the architectural ground truth, ensuring that synthesized radiance fields are spatially consistent with the physical acquisition environment. A top-down orthographic diagram of the experimental stage showing the registered positions of luminaires and structural elements in a shared coordinate system.
Figure S9 . Software-Defined Spatial Manifold: Front and Side Views. Registered orthographic projections map fixture elevations and lateral offsets, providing ground truth for spatial registration. Front and side orthographic diagrams of the experimental stage showing fixture elevations and lateral offsets used for spatial registration.
Figure S10 . Example of a GrandMA2 Console. The professional console provides robust stage-lighting control. A photograph of a GrandMA2 professional lighting console with multiple faders, displays, and control buttons.
Figure S11 . GrandMA2 Software Interface. The software provides real-time visualization, cue management, and timeline editing. A screenshot of the GrandMA2 software interface showing a stage preview, cue controls, and timeline editing panels.
Figure S12 . Timecode Encoding in GrandMA2. The interface precisely synchronizes multimodal lighting triggers; each marker denotes a discrete cue execution. A screenshot of a GrandMA2 timecode timeline with markers indicating discrete lighting-cue executions.
Figure S13 . Structured Parameter Space for Cue Sequences. (Left) The Sequence Pool stores illumination trajectories. (Right) An individual Cue Sequence specifies fixture behavior through an attribute vector that includes transition duration (Fade) and temporal offset (Delay), enabling deterministic execution. A two-part screenshot showing a cue-sequence repository on the left and detailed fixture-behavior parameters for one selected cue sequence on the right.