An important topic in the community of remote sensing is how to combine the complementary information provided by the huge amount of unlabeled multi-sensor data, such as Synthetic Aperture Radar (SAR) and optical images. Recently, contrastive learning methods have reached remarkable success in obtaining meaningful feature representations from multi-view data. However, these methods only focus on the image-level features, which may not satisfy the requirement for dense prediction tasks such as land-cover mapping. In this work, we propose a new self-supervised approach for SAR-optical data-fusion that can learn disentangled pixel-wise feature representations directly by taking advantage of both multi-view contrastive learning and the BYOL. The two key contributions were proposed for this approach: multi-view contrastive loss to encode the multi-modal images and the shift operation to reconstruct learned representations for each pixel by building the local consistency between different augmented views. With the aim to validate the effectiveness of the proposed approach, we conduct experiments on the land cover mapping task, where we trained the proposed approach using unlabeled SAR-optical image pairs while labeled data pairs were used for the linear classification and finetuning evaluations. We empirically show that the presented approach outperforms the state-of-the-art methods. In particular, it achieves an improvement on both linear classification and finetuning evaluations and reduces the dimension of representations with respect to the image-level contrastive learning method. Moreover, the proposed method is also validated to bring a sharp improvement on SAR-optical feature fusion than the early fusion fashion for the land-cover mapping task.