cs.CVMay 19, 2026

White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation

Authors: Shuwei Li, Lei Tan, Robby T. Tan

Organizations: National University of Singapore · ASUS Intelligent Cloud Services

Abstract

Color constancy aims to keep object colors consistent under varying illumination. Cross-camera generalization in color constancy remains challenging because learning-based models often overfit to the color response characteristics of the training camera, resulting in degraded performance on images captured by other cameras. We propose VLM-CC, a feedback-guided framework that formulates color constancy as an iterative refinement process. Instead of directly estimating the illuminant from raw input, VLM-CC performs iterative correction driven by vision-language model (VLM)-based evaluation. At each iteration, the image is white-balanced using the current estimate and converted to pseudo-sRGB. A lightweight LoRA-tuned VLM then assesses the corrected image, identifying the dominant residual color cast and providing qualitative feedback. This feedback is mapped to a residual illumination direction (red, green, or blue) and used to update the illuminant estimate until convergence. Our key idea is to reframe color constancy as an iterative perceptual feedback problem, leveraging VLM evaluation instead of direct RGB regression. By replacing direct RGB estimation with VLM-guided perceptual feedback, VLM-CC achieves state-of-the-art robustness in cross-camera color constancy across multiple datasets. Code will be available at https://github.com/NothingIknow/VLM-CC.

Explore similar work

Jan 8, 2026cs.CV

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes

Nighttime color constancy still remains a challenging problem in computational photography due to low-light noise and complex illumination conditions. We present RL-AWB, a novel framework combining statistical methods with deep reinforcement learning for nighttime white balance. Our method begins with a statistical algorithm tailored for nighttime scenes, integrating salient gray pixel detection with novel illuminant estimation. Building on this foundation, we develop the first deep reinforcement learning approach for color constancy that leverages the statistical algorithm as its core, mimicking professional AWB tuning experts by dynamically determining image-specific parameters at inference time, without requiring ground-truth illuminants or reference images. To further facilitate cross-sensor evaluation, we introduce the first multi-sensor nighttime dataset. Experiment results demonstrate that our method achieves strong generalization capability across low-light and well-illuminated images. Project page: https://ntuneillee.github.io/research/rl-awb/
Yuan-Kang Lee, Kuan-Lin Chen, Chia-Che Chang +1
Aug 7, 2026cs.RO

Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies

Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, depth, or point-cloud inputs. The wrist stream is masked throughout to prevent an unperturbed visual shortcut from confounding attribution to scene-camera variation. For flow-based VLAs, we propose to regularize the action-flow velocity field, the quantity directly integrated to generate continuous action chunks. We construct action-equivalent view pairs by resetting original LIBERO demonstrations to the same MuJoCo state and rendering nominal and perturbed scene-camera views. Both views are supervised by flow matching, while a cross-view loss encourages their predicted action-flow velocities to agree at the same sampled flow coordinates. On the LIBERO-Plus camera-perturbation track, our method reaches 87.2±\pm0.4% (4,797 rollouts per seed across 3 training seeds), +7.4pp over flow-matching-only training on the same paired data (79.8±\pm0.8%, also 3 seeds) and +12.5pp over naive mixed-camera SFT, while maintaining nominal-camera ID performance (95.0±\pm0.8%; same-data FM-only: 95.0±\pm4.3%). A shuffled-pair control collapses to 25.8%, showing that the gain depends on action-equivalent pairing. On a real robot, we evaluate three tabletop tasks with 10 rollouts per task and camera placement; held-out-camera success improves from 53.3% to 74.4% under the same single-scene-RGB inference interface.
Bingqi Huang, Bingchuan Wei, Xuan Wang +2
Jul 13, 2026cs.CV

Illuminant-Adaptive 3D Lookup Tables for Camera Color Correction

Color correction is a key component of camera image signal processing (ISP) pipelines, encompassing illuminant discounting and colorimetric mapping of device-dependent sensor responses to device-independent color spaces, such as CIE XYZ. Despite extensive research, accurate color correction remains challenging due to the non-linear relationship between camera sensor responses and CIE XYZ color space, as well as to the increasing presence of highly chromatic and spectrally complex LED illuminants. We propose a color correction framework based on illuminant-adaptive three-dimensional lookup tables (LUTs), which we call Color Correction LUT (C2^2LUT). Our method combines a chromaticity-aware illuminant representation with a non-linear color transformation, enabling accurate correction under illuminants spanning a wide range of chromaticities and spectral complexities. We employ Tucker tensor decomposition to represent the LUTs, ensuring that computational requirements remain sufficiently low for deployment in camera ISPs. In addition, we introduce a large-scale illuminants dataset comprising 1,473 spectral power distributions, with different chromaticities and spectral profiles. Experiments across multiple cameras, illuminants, reflectance datasets, and real captured images demonstrate consistent improvements over existing methods for color correction, reducing CIE ΔE00ΔE_{00} by up to 20% and angular error by up to 18% while remaining compatible with modern camera hardware constraints. Code and datasets are available at https://github.com/claudiom4sir/C2LUT.
Claudio Rota, Luca Cogo, Simone Bianco +1