According to various implementations, a method is performed at an electronic device with one or more processors, a non-transitory memory, a display. The electronic device optionally also includes an image sensor and input device(s). The method includes obtaining a semantic value that is associated with a physical object, based on image data. The image data is associated with a first input modality. In some implementations, the method includes obtaining a widget based on the semantic value, and displaying the widget according to an object-proximity criterion (e.g., display-locked, body-locked, or world-locked) with respect to the physical object. In some implementations, the method includes obtaining user data from the input device(s). The user data is associated with a second input modality that is different from the first input modality. Moreover, the method includes selecting a widget based on the semantic value and the user data, and displaying the widget.
Full Text
What is claimed is: