Scinovex
article Open AccessTop 1% cited

Neural volumes

ACM Transactions on Graphics · 2019 · Vol. 38(4) · pp. 1–14
Stephen LombardiTomas SimonJason SaragihGabriel SchwartzAndreas LehrmannYaser Sheikh

Abstract

Modeling and rendering of dynamic scenes is challenging, as natural scenes often contain complex phenomena such as thin structures, evolving topology, translucency, scattering, occlusion, and biological motion. Mesh-based reconstruction and tracking often fail in these cases, and other approaches (e.g., light field video) typically rely on constrained viewing conditions, which limit interactivity. We circumvent these difficulties by presenting a learning-based approach to representing dynamic objects inspired by the integral projection model used in tomographic imaging. The approach is supervised directly from 2D images in a multi-view capture setting and does not require explicit reconstruction or tracking of the object. Our method has two primary components: an encoder-decoder network that transforms input images into a 3D volume representation, and a differentiable ray-marching operation that enables end-to-end training. By virtue of its 3D representation, our construction extrapolates better to novel viewpoints compared to screen-space rendering techniques. The encoder-decoder architecture learns a latent representation of a dynamic scene that enables us to produce novel content sequences not seen during training. To overcome memory limitations of voxel-based representations, we learn a dynamic irregular grid structure implemented with a warp field during ray-marching. This structure greatly improves the apparent resolution and reduces grid-like artifacts and jagged motion. Finally, we demonstrate how to incorporate surface-based representations into our volumetric-learning framework for applications where the highest resolution is required, using facial performance capture as a case in point.

Advanced Vision and ImagingComputer Graphics and Visualization TechniquesRobotics and Sensor-Based LocalizationComputer scienceRendering (computer graphics)Artificial intelligenceComputer visionGridComputer graphics (images)Mathematics
Citations
647
FWCI
26.37
field-weighted impact
References
70
Percentile
100%
vs. same field & year
Citations per year
References
Real-time 3D reconstruction at scale using voxel hashing
ACM Transactions on Graphics · 2013 · 989 citations
A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms
International Journal of Computer Vision · 2002 · 6,694 citations
Accurate, Dense, and Robust Multiview Stereopsis
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2009 · 2,987 citations
A Theory of Shape by Space Carving
International Journal of Computer Vision · 2000 · 1,250 citations
Learning-based view synthesis for light field cameras
ACM Transactions on Graphics · 2016 · 696 citations
Deep video portraits
ACM Transactions on Graphics · 2018 · 649 citations
Stereo magnification
ACM Transactions on Graphics · 2018 · 708 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.