Abstract
We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos.
Monocular reconstruction is inherently ill-posed due to the absence of sufficient multiview constraints,
making accurate recovery of object geometry and temporal coherence particularly challenging.
To address this, we leverage multiple geometric cues, such as depth, optical flow, dynamic object segmentation, and tracking trajectories,
combined with rigidity constraints to estimate preliminary 3D scene dynamics during an advanced initialization stage.
Recovering scene flow prior to the photometric optimization phase reduces reliance on motion inference from visual appearance alone, which is often ambiguous in monocular settings.
To enable compact representations, fast training, and real-time rendering while supporting non-rigid deformations, the scene is decomposed into static and dynamic components,
with dynamic trajectories represented as time-dependent Poly-Fourier curves for parameter-efficient motion encoding.
We demonstrate that MOSAIC-GS achieves substantially faster optimization and rendering compared to existing methods,
while maintaining reconstruction quality on par with state-of-the-art approaches across standard monocular dynamic scene benchmarks.