NeurIPS 2026

Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling

Juntong Li*Lingwei Dang*†Haomin WuZiyan QiuQingxin XiaoQingyao Wu‡

School of Software Engineering, South China University of Technology

* Equal contributions† Project leader‡ Corresponding author

A sharper eye for detail.
The same inference architecture.

Align local geometry and global semantics to unlock fine-grained visual understanding.

PCA feature maps of buildings, cats, and a car compare CLIP, DINOv2, un2CLIP, KUEA, and SALM, alongside a radar plot of downstream performance.
Better structure. Broader capabilities. SALM produces spatially coherent features and improves zero-shot recognition, dense prediction, and multimodal understanding.
Image-only alignmentNo text supervision for alignmentNo extra inference cost

The idea

Abstract

Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP’s visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space.

To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring external textual supervision. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space.

Furthermore, driven by the empirical observations that CLIP’s shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP’s intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP’s zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs.

How it works

Local structure meets global semantics.

Two complementary paths enhance visual representations while preserving CLIP’s image-text alignment.

SALM architecture showing trainable CLIP, frozen CLIP and DINO encoders, dual-matrix alignment, and masked reconstruction through a Cross-Guided Adapter.
The SALM framework. Dual-Matrix Alignment transfers local geometry; latent masked modeling integrates fine-grained semantics; reference regularization preserves the original multimodal space.
01 / EXPLICIT

Align the structure

Dual-Matrix Alignment matches patch-to-patch similarities through a Spatial Relation Matrix and relative feature magnitudes through an Energy Difference Matrix.

02 / IMPLICIT

Recover the details

A Cross-Guided Adapter reconstructs masked reference features from visible context and CLIP features. A 75% masking ratio encourages the recovery of fine-grained semantics.

03 / PRESERVE

Keep the alignment

Regularization against the frozen original CLIP encoder protects image-text alignment. At inference, only the fine-tuned CLIP encoder is retained.

SALM-Self

The teacher is already inside.

CLIP’s shallow layers already capture rich spatial details. SALM-Self uses these features as its reference, unlocking fine-grained perception without an external vision model.

A teacher updated by exponential moving average (EMA) provides shallow features. Instead of a reconstruction loss, dual-matrix alignment compares reconstructed shallow features with the teacher’s features, preserving structure without forcing the student to reproduce low-level noise.

No external teacher model required
SALM-Self uses an EMA CLIP teacher, masked shallow features, and dual-matrix alignment for self-distillation.
SALM-Self architecture

Experiments

Fine-grained gains. Across tasks.

SALM aligns CLIP ViT-L/14-336 with DINOv2 ViT-L/14 with registers using only ImageNet-1K images. SALM-Self uses CLIP’s own shallow features.

Zero-shot recognition

69.25%

+1.90 pp vs. CLIP

11-dataset average · Table 1

Semantic segmentation

52.07 mIoU

+6.27 mIoU vs. CLIP

5-dataset mean · Linear probing · Table 2

Multimodal understanding

66.69%

+2.04 pp vs. LLaVA-1.5-7B + FT

8-benchmark average · Table 4

Results shown are for SALM. The zero-shot average excludes ImageNet-1K. Segmentation reports the arithmetic mean across ADE20K, Cityscapes, VOC2012, COCO-Stuff, and PASCAL Context, calculated from Table 2. Multimodal evaluation uses LLaVA-1.5-7B with LoRA instruction tuning on the same SFT data; image-only training refers to the encoder alignment stage. “pp” denotes percentage points.

A closer look

See what the features capture.

Clearer object boundaries, more coherent regions, and more discriminative representations.

PCA comparisons, from left to right: input image, CLIP, DINOv2, un2CLIP, KUEA, SALM-Self, and SALM.
Fine-grained local features. SALM captures semantic nuances and preserves local detail across diverse scenes.
t-SNE plots, from left to right: CLIP, KUEA, un2CLIP, and SALM, with SALM showing more compact and separated semantic clusters.
Discriminative global semantics. t-SNE visualizations show more compact and well-separated feature clusters.

Reference

BibTeX

Scroll to explore the figure · Press Esc to close