Ongoing
SEPIA Spatial Egocentric Perception for Interactive Agentsactive
A perception stack that lets interactive agents build and query spatial memory from first-person video — grounding what an agent sees into where it can act.
Show2Instruct Vision-grounded natural-language interfacesactive
Combining computer vision and large language models so natural-language commands can reference objects detected in the surrounding environment — demonstrated in construction, where a site inspection can ask in plain language whether a window or door meets BIM specifications and accessibility standards.
Dynamo CORE Labsactive
Dynamic scene understanding for autonomy, developed within CORE Labs — reasoning about motion and change in unstructured environments.
PRiSM Perception, reasoning & spatial modellingexploratory
Early-stage work connecting multimodal reasoning to explicit spatial models for robot decision-making.