HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models
Abstract
Vision-language models (VLMs) achieve strong cross-modal alignment but remain brittle to negation, often relying on shallow word associations rather than compositional reasoning. Fine-tuning on negation-specific data can also compromise their general purpose capabilities through catastrophic forgetting. We introduce HANCLIP (Hyperbolic, Angular, and Negation), a geometry-aware framework that improves negation sensitivity while preserving the structure of the pretrained joint embedding space. HANCLIP combines a hyperbolic contrastive objective, which models hierarchical relations and semantic asymmetries, with an angular triplet loss that separates negated descriptions from their affirmative counterparts. Using only 20,000 image-text quadruplets, HANCLIP consistently improves performance across CLIP, LongCLIP, and SmartCLIP backbones on the NegBench benchmark, while maintaining or improving zero-shot classification and image-text retrieval performance. These results show that lightweight, geometry-guided objectives can enhance negation understanding without large-scale retraining.