Summarize and Search: Learning Consensus-aware Dynamic Convolution for Co-Saliency Detection

by   Ni Zhang, et al.

Humans perform co-saliency detection by first summarizing the consensus knowledge in the whole group and then searching corresponding objects in each image. Previous methods usually lack robustness, scalability, or stability for the first process and simply fuse consensus features with image features for the second process. In this paper, we propose a novel consensus-aware dynamic convolution model to explicitly and effectively perform the "summarize and search" process. To summarize consensus image features, we first summarize robust features for every single image using an effective pooling method and then aggregate cross-image consensus cues via the self-attention mechanism. By doing this, our model meets the scalability and stability requirements. Next, we generate dynamic kernels from consensus features to encode the summarized consensus knowledge. Two kinds of kernels are generated in a supplementary way to summarize fine-grained image-specific consensus object cues and the coarse group-wise common knowledge, respectively. Then, we can effectively perform object searching by employing dynamic convolution at multiple scales. Besides, a novel and effective data synthesis method is also proposed to train our network. Experimental results on four benchmark datasets verify the effectiveness of our proposed method. Our code and saliency maps are available at <>.


page 3

page 6

page 7

page 8


Gradient-Induced Co-Saliency Detection

Co-saliency detection (Co-SOD) aims to segment the common salient foregr...

Discriminative Co-Saliency and Background Mining Transformer for Co-Salient Object Detection

Most previous co-salient object detection works mainly focus on extracti...

SESS: Saliency Enhancing with Scaling and Sliding

High-quality saliency maps are essential in several machine learning app...

DS-Net: Dynamic Spatiotemporal Network for Video Salient Object Detection

As moving objects always draw more attention of human eyes, the temporal...

CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds

We present a novel two-stage fully sparse convolutional 3D object detect...

AtLoc: Attention Guided Camera Localization

Deep learning has achieved impressive results in camera localization, bu...

Out of Sight, Out of Mind: A Source-View-Wise Feature Aggregation for Multi-View Image-Based Rendering

To estimate the volume density and color of a 3D point in the multi-view...

Please sign up or login with your details

Forgot password? Click here to reset