Visual Grounding
Visual grounding is the task of connecting natural language descriptions to corresponding regions within an image or 3D scene. Current research focuses on improving the accuracy and efficiency of visual grounding models, often employing transformer-based architectures and leveraging large multimodal language models (MLLMs) for enhanced feature fusion and reasoning capabilities. This field is crucial for advancing embodied AI, enabling robots and other agents to understand and interact with the world through natural language, and has significant implications for applications such as robotic manipulation, visual question answering, and medical image analysis.
Papers
September 3, 2023
August 24, 2023
August 23, 2023
August 22, 2023
August 21, 2023
August 18, 2023
August 8, 2023
July 25, 2023
July 23, 2023
July 21, 2023
July 18, 2023
July 17, 2023
July 12, 2023
June 13, 2023
June 10, 2023
June 6, 2023
May 24, 2023
May 23, 2023