AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance
Authors: Junyu Deng, Jiale Cao, Mengtian Li, Zhongxia Ji, Ruhua Chen, Yiyi He, Guangnan Ye, Zuo Hu
Organizations: Fudan University, Shanghai, China · Shanghai University, Shanghai, China · Shanghai Film Academy and Shanghai University, Shanghai, China · Shanghai Theatre Academy, Shanghai, China · Nanjing University of the Arts, Nanjing, China
We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling. Lighting design in live performance settings requires the seamless translation of musical features into dynamic lighting behaviors. However, traditional workflows remain time-consuming, labor-intensive, and difficult to transfer. AuraLuxMuse encodes music and professional cue sequences into a shared retrieval space, estimates cue-event density, and retargets selected fixture commands to the destination stage. It assists pre-production authoring by returning editable cues rather than replacing the designer with an unconstrained generator. At the heart of AuraLuxMuse are two key modules: Lighting-Aligned Music Pretraining (LAMP), which performs contrastive learning between audio and lighting cues for alignment, and Preference-Adaptive Mixture of Experts (PAMoE), which conditions preference-aware cue retrieval and adaptation on designers' intent through a gated ensemble of style-specific expert networks. To support training and evaluation, we introduce Musilux, the first dataset of paired musical audio and professional lighting cue sequences under diverse performance scenarios. We evaluate AuraLuxMuse across both virtual simulation environments and professional-grade laboratories. Experimental results, including objective and subjective evaluation, demonstrate that AuraLuxMuse retrieves and adapts stage-lighting cues that are visually cohesive, semantically meaningful, and artistically expressive, showing its potential for AI-assisted aesthetic stage design.
Figures & tables
Figure 1. Procedure Comparison: Traditional Stage Lighting Design vs. AuraLuxMuse. Conventional pipelines require designers to manually decompose music into artistic components, then iteratively design cues, align timing, and input parameters into consoles. AuraLuxMuse accelerates this workflow through feature extraction, preference-conditioned retrieval, cue adaptation, and console dispatch. Comparison of a traditional manual stage-lighting workflow with the AuraLuxMuse workflow, from music input and artistic analysis to console-ready lighting output.
Figure 2. Overview of the AuraLuxMuse Pipeline. Music and lighting data are individually preprocessed before entering LAMP for contrastive learning, and preference representations are injected via PAMoE. P1/P2/P3 denote per-style frame counts, and Ti=∑t=1i∣Pt∣ is the cumulative sample count used to index the illustrated segments. Cosine similarity and the metadata value K predicted by the non-causal TCN Metadata Predictor are used to query the cue corpus. AuraLuxMuse pipeline showing music and cue encoders, LAMP contrastive alignment, PAMoE preference routing, the non-causal TCN Metadata Predictor, top-K cue retrieval, and console execution.
Figure 3. Musilux Construction Procedure. Music and lighting cues are collected, aligned both within cue structures and across modalities, and then aggregated into the Cue Sequence Corpus that constitutes the dataset. Flow diagram of the Musilux dataset construction procedure: collection, intra-cue alignment, cross-modal alignment, and aggregation into the Cue Sequence Corpus.
Figure 4. Qualitative Results. AuraLuxMuse and the manual design exhibit different spatial and temporal choices in the same Drama scenario; these examples illustrate scenario-dependent competitiveness rather than universal superiority. Side-by-side qualitative comparison frames between manually designed lighting and AuraLuxMuse-generated lighting in a Drama scenario, illustrating temporal continuity.
Figure 5. Stage Demonstrations across Lighting Styles. Lighting configurations are visualized across varying levels of Color harmony, Brightness control, and Zone distribution. Definitions of stage-lighting terms are provided in parentheses. Grid of stage demonstrations showing different combinations of color harmony, brightness levels, and zone distributions for the lighting style space.
Scenario
Source
FID
Sexp↑
Snon
EGSS ↑
Concert
Manual
8.21
2.58 (1.33)
4.31 (0.73)
3.09 (1.15)
AuraLuxMuse
3.65 (1.00)
3.69 (0.88)
3.66 (0.75)
Drama
Manual
4.47
3.08 (1.10)
4.03 (0.93)
3.36 (1.04)
AuraLuxMuse
2.71 (1.24)
3.27 (1.11)
2.88 (0.93)
Entertainment
Manual
2.00
2.46 (1.25)
3.31 (0.98)
2.71 (1.16)
AuraLuxMuse
3.21 (1.25)
3.39 (1.13)
3.26 (0.94)
Table 1. Quantitative Results. Average expert and non-expert ratings are reported for all scenarios. AuraLuxMuse achieves scenario-dependent competitiveness with the manual counterparts. Values in parentheses are standard deviations. FID measures the representational discrepancy between AuraLuxMuse and the corresponding reference.
Strategy
FID ↓
Precision@5 ↑
mAP ↑
EGSS ↑
Random
6.32
0.07
0.03
2.71 (0.72)
Position-based
25.14
0.07
0.04
1.98 (0.73)
Frequency-based
16.32
0.07
0.03
2.44 (0.76)
Nearest-neighbor
21.65
0.08
0.04
2.10 (0.73)
Rule-based
6.61
0.06
0.02
2.81 (0.81)
Similarity-based
4.89
0.17
0.10
3.27 (0.88)
Table 2. Comparison of Retrieval Strategies. We compare similarity retrieval with random, position-based, frequency-based, nearest-neighbor, and rule-based strategies. Values in parentheses are standard deviations.
Table 3. Subjective Results across Music Styles. EGSS is reported for six validation styles and six test styles. Validation uses held-out examples with manual references; test tracks and the test stage are excluded from training.
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1 . Qualitative Results for AuraLuxMuse Variants in the Drama Scenario. Three architecture variants produce distinct retrieved/adapted lighting outputs as editable references for professional designers. A grid of stage-lighting frames comparing manual design with three AuraLuxMuse variants in the Drama scenario.
Figure S2 . Qualitative Results in the Concert Scenario. Consecutive frames compare manual lighting with AuraLuxMuse retrieved/adapted outputs under the same scenario; the variants show how the model structure changes cue selection. A grid of stage-lighting frames comparing manual design with AuraLuxMuse variants in the Concert scenario, together with music features and event timestamps.
Figure S3 . Qualitative Results in the Entertainment Scenario. Consecutive frames compare a manual design with AuraLuxMuse retrieved/adapted cues under the same Entertainment scenario. The example illustrates different choices in visual complexity and atmosphere without implying universal superiority. A grid of consecutive Entertainment-scenario lighting frames comparing manual design with AuraLuxMuse retrieved and adapted cues, plus music features and event timestamps.
Ratings
Color Harmony
1
Harsh, unbalanced colors; poor harmony
2
Colors clash; mildly unbalanced
3
Colors blend acceptably; moderate harmony
4
Colors blend smoothly; strong harmony
5
Colors blend flawlessly; outstanding harmony
Appendix
Table S1 . Lighting Evaluation for Experts: Color Harmony. This dimension rates the aesthetic balance and blending of stage-lighting colors across the complete visual composition of each scene.
Ratings
Brightness Control
1
Uneven brightness; excessively dim/bright
2
Uneven brightness; slightly harsh
3
Reasonable brightness control; minor issues
4
Consistent brightness; well-adjusted
5
Perfect brightness; optimal balance
Appendix
Table S2 . Lighting Evaluation for Experts: Brightness Control. This dimension rates the consistency and adjustment of brightness across an entire performance sequence under review.
Ratings
Zone Distribution
1
Uneven zones; inadequate coverage
2
Uneven zones; noticeable gaps
3
Fair zoning; slight inconsistencies
4
Even zoning; effective coverage
5
Flawless zoning; complete coverage
Appendix
Table S3 . Lighting Evaluation for Experts: Zone Distribution. This dimension rates evenness and coverage across the visible stage area.
Ratings
Rhythmic Tempo
1
No rhythm; disorganized timing
2
Weak rhythm; moderately off
3
Average rhythm; reasonably aligned
4
Good rhythm; accurately timed
5
Excellent rhythm; impeccably timed
Appendix
Table S4 . Lighting Evaluation for Experts: Rhythmic Tempo. This dimension rates lighting synchronization with performance rhythm.
Ratings
Narrativity
1
Fails to convey plot / emotion
2
Weakly conveys plot / emotion
3
Moderately conveys plot / emotion
4
Clearly conveys plot / emotion
5
Strongly conveys plot / emotion
Appendix
Table S5 . Lighting Evaluation for Experts: Narrativity. This dimension rates how effectively lighting conveys plot and emotion.
Ratings
Evaluation Dimensions
Rhythmic Alignment
Thematic Suitability
Visual Consistency
1
No alignment
Not suited
Extremely inconsistent
2
Poor alignment
Poorly suited
Somewhat inconsistent
3
Basic alignment
Basically suited
Basically consistent
4
Fair alignment
Fairly suited
Generally consistent
5
Perfect alignment
Perfectly suited
Always consistent
Appendix
Table S6 . Lighting Evaluation Dimensions (Relative to Music) for Non-Experts. The criteria capture spectators’ intuitive perceptions of performance quality across the three evaluated scenarios.
Peripheral, Atmospheric, Neutral, Attention, Dominant Zone
Rhythmic Tempo
Static, Slow, Moderate, Fast, Rapid
Appendix
Table S7 . Overview of Lighting Style Dimensions. Color, Brightness, Zone, and Rhythm are mapped to representative attributes.
Level
Description
Illuminance (lux)
Dark
Almost invisible
0–20
Dim
Barely visible
20–100
Moderate
Visually comfortable
100–300
Bright
Clear visibility
300–700
Intense
Extremely bright
700+
Appendix
Table S8 . Brightness Attribute Levels. Each level is defined by a perceptual description and illuminance range for preference specification.
Level
Description
Peripheral
Invisible boundary or background layer
Atmospheric
Subtle lighting for mood creation
Neutral
Balanced fill without drawing attention
Attention
Guides audience attention
Dominant
Strongest visual focus
Appendix
Table S9 . Zone Distribution Attribute Levels. Each level defines a degree of visual prominence within a scene.
Level
Description
Static
Minimal change, nearly static
Slow
Gentle & Sparse transitions, rare rhythmic cues
Moderate
Regular tempo, structured but not fast
Fast
Frequent changes, strong rhythmic presence
Rapid
Intense flashes with strong beats
Appendix
Table S10 . Rhythmic Tempo Attribute Levels. Each level defines a transition pace for preference specification.
Figure S4 . Patching Mechanism Overview. Retrieved artistic cues are retargeted to compatible fixtures through a logical layout map and an industrial-console patch table used for dispatch to the console. A schematic showing retrieved lighting cues being retargeted through compatibility checks, a logical layout map, and a patch table before console control.
Music Encoder / Cue Encoder
Transformer
ResNet
VGGish + MLP
0.699( ↓ )
2.015
VGGish + ResNet
2.181
3.926
Appendix
Table S11 . Ablation of LAMP Architectures. Contrastive loss is reported after 500 training epochs of model optimization.
Model Variant
FID (↓)
Sexp
Snon
EGSS (↑)
PAMoE Only
18.06
2.26
2.86
2.44
Soft Only
10.97
2.83
3.47
3.02
PAMoE + Soft
4.89
3.19
3.45
3.27
Appendix
Table S12 . Ablation of PAMoE and Soft Alignment. All experiments use the same learning rate, weight decay, and number of epochs.
Architecture
MAE ↓
RMSE ↓
R 2 ↑
Pearson r↑
Bias ↓
MLP
16.987
24.113
-368.715
0.061
+16.973
TCN Density
0.343
0.449
0.872
0.934
+0.017
Appendix
Table S13 . Metadata Predictor Performance. We compare window-level cue counting with a baseline MLP and our non-causal TCN.
Architecture
FID ↓
EGSS ↑
Precision@5 ↑
mAP ↑
CLAP
5.06
2.52
0.11
0.043
SLAP
7.20
2.11
0.02
0.085
AudioCLIP
5.21
2.18
0.13
0.05
AuraLuxMuse
4.89
3.27
0.170
0.100
Appendix
Table S14 . Comparison with Contrastive Baselines. AuraLuxMuse obtains the highest Precision@5 and EGSS under the reported setup.
Architecture
FID
EGSS ↑
Inference Time (s/file) ↓
T5
4.28
1.79
25.74
BART
4.21
1.89
35.20
Autoformer
3.98
1.24
207.27
SpecTNT
3.95
1.24
150.18
AuraLuxMuse
4.89
3.27
7.65
Appendix
Table S15 . Comparison with Generative Baselines. Against four sequence-to-sequence models, AuraLuxMuse achieves lower inference latency and a higher EGSS under the reported setup.
Strategy
K
FID ( ↓ )
EGSS ( ↑ )
Fixed Small
1
8.42
2.85
Fixed Large
10
9.82
3.02
Adaptive
Dynamic
4.89
3.27
Appendix
Table S16 . Ablation Study: Static vs. Dynamic K Values. Dynamic K yields a lower FID and a higher EGSS than either static setting.
Soft Weight ( w )
Avg FID ↓
R@1 ↑
R@5 ↑
R@10 ↑
MRR ↑
nDCG ↑
w=0.1
11.85
0.045
0.223
0.338
0.144
0.303
w=0.3
4.89
0.046
0.249
0.374
0.149
0.309
w=0.5
10.10
0.062
0.202
0.280
0.143
0.300
w=0.7
61.21
0.033
0.159
0.301
0.118
0.279
w=0.9
47.09
0.020
0.127
0.258
0.097
0.261
Appendix
Table S17 . Ablation on Soft Alignment Weight ( w ). We evaluate the impact of the soft-alignment weight w on cross-modal retrieval performance and Fréchet Inception Distance (FID). The hard-alignment weight is fixed at 1.0. Results indicate that w=0.3 yields the best balance, achieving the strongest overall retrieval metrics by leveraging stylistic similarities without blurring discriminative boundaries.
Embedding
FID (↓)
Sexp
Snon
EGSS (↑)
One Hot
12.24
2.51
3.29
2.74
Attribute Map
4.89
3.19
3.45
3.27
Appendix
Table S18 . Ablation of Cue-Embedding Strategies. The architecture is fixed, and only the embedding strategy varies.
Figure S5 . Sensitivity analysis of the EGSS metric with respect to λ . The parameter λ controls the weight assigned to expert preferences in the subjective evaluation process. The manual lighting design serves as an approximate upper bound. AuraLuxMuse consistently outperforms all automated baselines across the full λ range, demonstrating strong robustness to this hyperparameter. The vertical dashed line indicates our selected default configuration ( λ=0.7 ). A line chart showing EGSS for manual design and automated baselines as the expert-preference weight lambda varies, with a dashed line marking lambda equal to 0.7.
Figure S6 . Examples of Stage Lighting Fixtures: (a-b) Imaging Lights: FINE 300C LEKO and FINE 400T/D LEKO for sharp projection and patterned illumination; (c-d) Spotlights: FINE 420 BEAM IP and FINE 480 BSW IP for dynamic focus and concentrated beams; (e-f) Soft Lights: FINE 600T/D PANEL and FINE 300T/D PANEL for diffused illumination and even coverage; (g-h) Floodlights: FINE 1514 DG and FINE 1514 ZOOM for broad lighting and wide effects. An eight-panel image sheet showing imaging lights, spotlights, soft lights, and floodlights used in the stage-lighting setup.
Figure S7 . Fixture Parameter Visualization. The Fixture Sheet records instantaneous lighting-infrastructure states. Its fields match the functional components in Section S4 , enabling precise state tracking during acquisition. A fixture-sheet table listing the instantaneous state parameters of the lighting infrastructure for data acquisition and tracking.
Figure S8 . Software-Defined Spatial Manifold: Top-down View. The top-down orthographic projection of the experimental environment as registered within the GrandMA2 configuration space. Each luminaire instance and structural element is precisely localized within a unified global coordinate system, providing the necessary extrinsic parameters for spatiotemporal alignment. This geometric blueprint serves as the architectural ground truth, ensuring that synthesized radiance fields are spatially consistent with the physical acquisition environment. A top-down orthographic diagram of the experimental stage showing the registered positions of luminaires and structural elements in a shared coordinate system.
Figure S9 . Software-Defined Spatial Manifold: Front and Side Views. Registered orthographic projections map fixture elevations and lateral offsets, providing ground truth for spatial registration. Front and side orthographic diagrams of the experimental stage showing fixture elevations and lateral offsets used for spatial registration.
Figure S10 . Example of a GrandMA2 Console. The professional console provides robust stage-lighting control. A photograph of a GrandMA2 professional lighting console with multiple faders, displays, and control buttons.
Figure S11 . GrandMA2 Software Interface. The software provides real-time visualization, cue management, and timeline editing. A screenshot of the GrandMA2 software interface showing a stage preview, cue controls, and timeline editing panels.
Figure S12 . Timecode Encoding in GrandMA2. The interface precisely synchronizes multimodal lighting triggers; each marker denotes a discrete cue execution. A screenshot of a GrandMA2 timecode timeline with markers indicating discrete lighting-cue executions.
Figure S13 . Structured Parameter Space for Cue Sequences. (Left) The Sequence Pool stores illumination trajectories. (Right) An individual Cue Sequence specifies fixture behavior through an attribute vector that includes transition duration (Fade) and temporal offset (Delay), enabling deterministic execution. A two-part screenshot showing a cue-sequence repository on the left and detailed fixture-behavior parameters for one selected cue sequence on the right.
Music-inspired Automatic Stage Lighting Control (ASLC) has gained increasing attention in recent years due to the substantial time and financial costs associated with hiring and training professional lighting engineers. However, existing methods suffer from several notable limitations: the low interpretability of rule-based approaches, the restriction to single-primary-light control in music-to-color-space methods, and the limited transferability of music-to-controlling-parameter frameworks. To address these gaps, we propose SeqLight, a hierarchical deep learning framework that maps music to multi-light Hue-Saturation-Value (HSV) space. Our approach first customizes SkipBART, an end-to-end single primary light generation model, to predict the full light color distribution for each frame, followed by hybrid Imitation Learning (IL) techniques to derive an effective decomposition strategy that distributes the global color distribution among individual lights. Notably, the light decomposition module can be trained under varying venue-specific lighting configurations using only mixed light data and no professional demonstrations, thereby flexibly adapting across diverse venues. In this stage, we formulate the light decomposition task as a Goal-Conditioned Markov Decision Process (GCMDP), construct an expert demonstration set inspired by Hindsight Experience Replay (HER), and introduce a three-phase IL training pipeline, achieving strong generalization capability. To validate our IL solution for the proposed GCMDP, we conduct a series of quantitative analysis and human study. The code and trained models are provided at https://github.com/RS2002/SeqLight .
Zijian Zhao, Dian Jin, Zijing Zhou +1
The Hong Kong University of Science and Technology · The Hong Kong Polytechnic University · The University of Hong Kong +1
Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, or mood progression. To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent. Scoring items individually makes evaluation diagnostic by intent source and musical dimension, rather than a single opaque score. We instantiate this as MuRA-Bench, a benchmark of real-world platform requests curated by music experts. We further propose MIRA (Musical Intent Refinement Agent), a test-time agent that first grounds a request's intent into rubrics, then searches over prompt revisions for a black-box generator under a bounded budget, iteratively generating music, verifying it against the rubrics, and using this feedback to guide a trajectory-aware tree search. Experiments across open-source and commercial backends show that MIRA improves intent alignment, enabling an open-source generator to achieve performance comparable to representative commercial systems (e.g. Suno and Mureka). Project page: https://mirareview.github.io/.
Zekai Liu, Zhilin Wang, Xuzheng He +2
Shandong University · University of Science and Technology of China · Central Conservatory of Music +2
Large audio language models (LALMs) have shown promising progress in broad music-understanding tasks such as tagging, retrieval, and captioning. Music understanding that requires finer hearing over both the content and how it is realized within a performance through dynamics, phrasing, articulation, time, and other performance techniques, however, remains at an earlier stage. Existing audio-language model (ALM) training pipelines typically rely on coarse, weakly grounded captions and therefore provide little support for learning these subtle nuances in music, limiting their ability to serve real-world applications in education or artistic practice. We therefore introduce MuNo-SP (Music Notation unifying Score and Performance), a text-based representation that jointly encodes score content and performance information. Building on MuNo-SP, we develop an automatic training-data generation pipeline that uses aligned scores and performances to produce long-form auditory analyses and musically informed question-answer pairs. We use this pipeline to construct MAESTROCaps, a classical piano dataset comprising 148 long-form performance analyses and 31,080 question-answer pairs derived from 148 aligned score-performance pairs. In a human evaluation, MuNo-SP analyses were preferred by majority vote over MIDI-only analyses for eight of nine excerpts. MuNo-SP also performed strongly on a benchmark of score-performance understanding, suggesting that integrating score and performance information enables more reliable and musically informative LALM supervision than a MIDI-only baseline.