Align the structure
Dual-Matrix Alignment matches patch-to-patch similarities through a Spatial Relation Matrix and relative feature magnitudes through an Energy Difference Matrix.
NeurIPS 2026
School of Software Engineering, South China University of Technology
* Equal contributions† Project leader‡ Corresponding author
The idea
Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP’s visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space.
To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring external textual supervision. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space.
Furthermore, driven by the empirical observations that CLIP’s shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP’s intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP’s zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs.
How it works
Two complementary paths enhance visual representations while preserving CLIP’s image-text alignment.
Dual-Matrix Alignment matches patch-to-patch similarities through a Spatial Relation Matrix and relative feature magnitudes through an Energy Difference Matrix.
A Cross-Guided Adapter reconstructs masked reference features from visible context and CLIP features. A 75% masking ratio encourages the recovery of fine-grained semantics.
Regularization against the frozen original CLIP encoder protects image-text alignment. At inference, only the fine-tuned CLIP encoder is retained.
SALM-Self
CLIP’s shallow layers already capture rich spatial details. SALM-Self uses these features as its reference, unlocking fine-grained perception without an external vision model.
A teacher updated by exponential moving average (EMA) provides shallow features. Instead of a reconstruction loss, dual-matrix alignment compares reconstructed shallow features with the teacher’s features, preserving structure without forcing the student to reproduce low-level noise.
No external teacher model required
Experiments
SALM aligns CLIP ViT-L/14-336 with DINOv2 ViT-L/14 with registers using only ImageNet-1K images. SALM-Self uses CLIP’s own shallow features.
Zero-shot recognition
69.25%+1.90 pp vs. CLIP
11-dataset average · Table 1Semantic segmentation
52.07 mIoU+6.27 mIoU vs. CLIP
5-dataset mean · Linear probing · Table 2Multimodal understanding
66.69%+2.04 pp vs. LLaVA-1.5-7B + FT
8-benchmark average · Table 4Results shown are for SALM. The zero-shot average excludes ImageNet-1K. Segmentation reports the arithmetic mean across ADE20K, Cityscapes, VOC2012, COCO-Stuff, and PASCAL Context, calculated from Table 2. Multimodal evaluation uses LLaVA-1.5-7B with LoRA instruction tuning on the same SFT data; image-only training refers to the encoder alignment stage. “pp” denotes percentage points.
A closer look
Clearer object boundaries, more coherent regions, and more discriminative representations.


Reference