ECCV 2026

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

A project page for our ECCV 2026 paper on grounding-aware remote sensing referring segmentation.

Tingzhang Luo Ruizhong Liu Yichao Liu Cheng Fan Yu Liu Jianyuan Guo
CROSS framework figure

Overview of CROSS. The framework integrates Linguistic-Guided Cascaded Distillation (LGCD) and Perspective-Spatial Contrastive Learning (PSCL) for robust remote sensing referring segmentation.

Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre-trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak-Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel-level structural guidance, causing localization drift; and (2) Object-Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS.

Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic-Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective-Spatial Contrastive Learning (PSCL) imposes cross-anchored constraints by mining mask-filtered deceptive distractors and spatial-linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency.

Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.

@inproceedings{luo2026cross,
  title={CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation},
  author={Luo, Tingzhang and Liu, Ruizhong and Liu, Yichao and Fan, Cheng and Liu, Yu and Guo, Jianyuan},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}