| ICLR area: applications to computer vision, audio, language, and other modalities |
| 0 (validated novel) | Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-world long-context information, primarily due to insufficient long-context alignment caused by data quality issues, training inefficiencies, and the lack of well-designed optimization objectives. To address these limitations, we propose a framework named Short-to-Long Preference Optimization (SoLoPO), decoupling long-context preference optimization (PO) into two components: short-context PO and short-to-long reward alignment (SoLo-RA), supported by both theoretical and empirical evidence. Specifically, short-context PO leverages preference pairs sampled from short contexts to enhance the model’s contextual knowledge utilization ability. Meanwhile, SoLo-RA explicitly encourages reward score consistency for the responses when conditioned on both short and long contexts that contain identical task-relevant information. This facilitates transferring the model’s ability to handle short contexts into long-context scenarios. SoLoPO is compatible with mainstream preference optimization algorithms, while substantially improving the efficiency of data construction and training processes. |
| 1 (weakly labeled lower-novelty) | Multimodal understanding requires more than aligning representations across vision, audio, and language—it demands reasoning about temporal precedence and causal relationships between modalities. Existing approaches treat modalities symmetrically through attention mechanisms, failing to capture how information in one modality temporally influences or explains events in another. We introduce Causal Multimodal Transformers, a framework that explicitly models directional causal dependencies across asynchronous modality streams. Our architecture incorporates learned temporal offsets and causal masking patterns that respect the natural information flow between modalities, enabling the model to distinguish between coincidental co-occurrence and genuine causal influence. By decomposing cross-modal attention into temporally-ordered causal graphs, the framework learns which modality provides predictive information for events in others at different time scales. This causally-aware design enhances interpretability and enables counterfactual reasoning—answering questions like “what would the visual scene be if this sound had not occurred?” Applications span video understanding, audio-visual speech processing, and multimodal content generation where temporal causality is fundamental to meaning. |
| Expert selected: 0 | The idea in 1 of causal/temporal dependence between modalities is old. Temporal offsets, causal masks with directional cross-modal attention etc. It’s all a generic shallow combination of these things. e.g., https://arxiv.org/abs/1906.00295 (>3K citations). The idea in 0 of decoupling long-context preference optimization into short and short-to-long seems clever and original, and a very specific method invention. |
| ICLR area: unsupervised, self-supervised, semi-supervised, and supervised representation learning |
| 0 (weakly labeled lower-novelty) | Representation learning paradigms—unsupervised, self-supervised, semi-supervised, and supervised—are typically treated as distinct methodologies requiring separate architectural designs and training protocols. We introduce Spectrum Learning, a unified framework that views supervision as a continuous spectrum rather than discrete categories. Our approach employs a meta-learned weighting mechanism that dynamically modulates the contribution of multiple learning objectives based on local supervision density in the feature space. By treating each data point as existing along a supervision gradient, the framework adaptively combines contrastive self-supervised losses, consistency regularization, pseudo-labeling, and supervised classification within a single coherent optimization. A novel supervision-aware attention module enables the model to identify which learning paradigm is most informative for different regions of the data manifold. This fluid integration allows seamless transitions as supervision availability changes, from fully unsupervised to fully supervised settings, without architectural modification. Spectrum Learning naturally handles mixed supervision scenarios common in practice, where different samples have varying levels of annotation quality and granularity, providing a principled approach to leveraging all available learning signals simultaneously. |
| 1 (validated novel) | Multi-view clustering integrates the consistency and complementarity of different views to achieve unsupervised data grouping. Existing multi-view clustering methods primarily confront two challenges: i) they generally perform feature extraction in the feature domain, which is sensitive to noise and may neglect cluster-specific information that is indistinguishable in the original space; ii) current dynamic fusion methods adopt static strategies to learn weights, lacking capability to adjust strategies adaptively under complex scenarios according to variations in data distribution and view quality. To address these issues, we propose a large language model assisted dynamic agent for multi-view clustering (LLM-DAMVC), a novel framework that recasts multi-view clustering as a dynamic decision-making problem orchestrated by a large language model. Specifically, each view is equipped with complementary agents dedicated to feature extraction. A dual-domain contrastive module is introduced to optimize feature consistency and enhance cluster separability in both the feature domain and frequency domain. Additionally, an LLM-assisted view fusion mechanism provides a flexible fusion weight learning strategy that can be adaptively applied to complex scenarios and significantly different views. |