Image Captioning

Momentum

7 papers in the last four weeks, down 12% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 59

All topics
CardsList
  1. Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning

    Jun 17, 2026Minh-Loi Nguyen, Xuan-Vu Le, Long-Bao Nguyen +2Multimodal RAGMultimodal IR

  2. CIAN: Multi-Stage Framework for Event-Enriched Image Captioning via Retrieval-Augmented Generation

    Jun 16, 2026Trinh Thi Thu Hien, Trung-Nghia LeMultimodal RAGImage Captioning

  3. Gaze Heads: How VLMs Look at What They Describe

    Jun 12, 2026Rohit Gandikota, David BauImage CaptioningVisual Attention

  4. Zero-Shot Captioning for Cultural Heritage: Automated Image Analysis of Traditional Indonesian Clothing

    Jun 11, 2026Anugrah Aidin Yotolembah, Novanto Yudistira, Gembong Edhi SetyawanRetrieval-Augmented GenerationImage Captioning

  5. CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

    Jun 8, 2026Penghui Yang, Long Xing, Xiaoyi Dong +10Image CaptioningMultimodal Pretraining

  6. Beyond Self-Attention: Sub-Quadratic Vision Transformers for Fast Image Captioning

    Jun 7, 2026Chiradeep Ghosh, Dakshina Ranjan KiskuImage CaptioningVision Transformer

  7. T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining

    May 30, 2026Tayeba Qazi, Ayush Maheshwari, Prerana Mukherjee +1Thermal ImagingImage Captioning

  8. A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models

    May 29, 2026Iosif Tsangko, Andreas Triantafyllopoulos, George Margetis +2Image CaptioningVLM Adaptation

  9. VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

    May 27, 2026Xingyu Lu, Jinpeng Wang, Yi-Fan Zhang +13Image CaptioningObject Hallucination in VLMs

  10. BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model

    May 20, 2026Gonçalo Gomes, Bruno Martins, Chrysoula ZervaVLM EvaluationImage Captioning

  11. Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments

    May 20, 2026Van Quang NguyenImage CaptioningEfficient ViTs

  12. ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

    May 19, 2026Tianle Li, Xuyang Shen, Yan Ma +7VLM EvaluationReward Modeling

  13. EmoMind: Decoding Affective Captions from Human Brain fMRI

    May 16, 2026Bilal A. Mohammed, Lin Gu, Ruogu FangImage CaptioningNeural Decoding

  14. HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

    May 14, 2026Man Wang, Chenyang Liu, Wenjun Li +5Remote Sensing Image UnderstandingImage Captioning

  15. Weather-Robust Scene Semantics with Vision-Aligned 4D Radar

    May 8, 2026Kali Hamilton, Christoffer HeckmanCross-Modal LearningVLM Robustness

  16. DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

    May 6, 2026Yuancheng Wei, Haojie Zhang, Linli Yao +7Image CaptioningMultimodal Model Evaluation

  17. Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

    May 4, 2026Lucrezia Tosato, Gianluca Lombardi, Ronny HanschRemote Sensing Image UnderstandingImage Captioning

  18. Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention

    Apr 30, 2026Nhi Ngoc-Yen Nguyen, Anh-Duc Nguyen, Nghia Hieu Nguyen +2Image CaptioningLow-Resource Language Processing

  19. CANVAS: Captioning Art with Narrative Visual-Audio AI Systems

    Apr 30, 2026Vignesh NagarajanImage Captioning

  20. Retrieval-Guided Generation for Safer Histopathology Image Captioning

    Apr 27, 2026Md. Enamul Hoq, Wataru Uegami, Saghir Alfasly +6Multimodal RAGComputational Pathology

  21. JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning

    Apr 27, 2026Swadhin Das, Vivek YadavRemote Sensing Image UnderstandingImage Captioning

  22. Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts

    Apr 20, 2026Run Xu, Lu Li, Rongzhao Zhang +1Image CaptioningMultimodal Generation

  23. Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models

    Dec 16, 2025George-Andrei Dima, Răzvan-Alexandru Smădu, Dumitru-Clementin CercelVisual Question AnsweringImage Captioning

  24. Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

    May 22, 2025Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad +1Image CaptioningMultimodal Model Evaluation