Computer Science > Computer Vision and Pattern Recognition

arXiv:2308.02191 (cs)

[Submitted on 4 Aug 2023]

Title:ES-MVSNet: Efficient Framework for End-to-end Self-supervised Multi-View Stereo

Authors:Qiang Zhou, Chaohui Yu, Jingliang Li, Yuang Liu, Jing Wang, Zhibin Wang

View PDF

Abstract:Compared to the multi-stage self-supervised multi-view stereo (MVS) method, the end-to-end (E2E) approach has received more attention due to its concise and efficient training pipeline. Recent E2E self-supervised MVS approaches have integrated third-party models (such as optical flow models, semantic segmentation models, NeRF models, etc.) to provide additional consistency constraints, which grows GPU memory consumption and complicates the model's structure and training pipeline. In this work, we propose an efficient framework for end-to-end self-supervised MVS, dubbed ES-MVSNet. To alleviate the high memory consumption of current E2E self-supervised MVS frameworks, we present a memory-efficient architecture that reduces memory usage by 43% without compromising model performance. Furthermore, with the novel design of asymmetric view selection policy and region-aware depth consistency, we achieve state-of-the-art performance among E2E self-supervised MVS methods, without relying on third-party models for additional consistency signals. Extensive experiments on DTU and Tanks&Temples benchmarks demonstrate that the proposed ES-MVSNet approach achieves state-of-the-art performance among E2E self-supervised MVS methods and competitive performance to many supervised and multi-stage self-supervised methods.

Comments:	arXiv admin note: text overlap with arXiv:2203.03949 by other authors
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2308.02191 [cs.CV]
	(or arXiv:2308.02191v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2308.02191

Submission history

From: Qiang Zhou [view email]
[v1] Fri, 4 Aug 2023 08:16:47 UTC (6,254 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:ES-MVSNet: Efficient Framework for End-to-end Self-supervised Multi-View Stereo

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:ES-MVSNet: Efficient Framework for End-to-end Self-supervised Multi-View Stereo

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators