A system generates an application specific pose estimation responsive to a camera pose, a scene pose and an object pose. The system includes a camera that generates an observed image of an object and a scene. A processor generates a current pose estimate having a current object pose and a current scene pose. The observed image is segmented using a 2D segmentation technique to obtain a first segmented component representing a scene image of the current scene pose and a second segmented component representing an object image representing the current object pose that is isolated from the first segmented component. An object render is generated by applying the object image to a first NeRF, a scene render is generated by applying the scene image to a second NeRF, and a final rendered output is generated by combining the object render and the scene render to generate a final pose estimate.
Full Text
What is claimed is: