Humans understand the physical world through an active process. When sensory evidence is insufficient, we interact: we lift an object to feel its weight, or shake it to hear what is inside. However, existing multi-sensory robot systems mostly integrate the inputs they are given, rather than actively acquire the evidence they are missing.
ROMA is an LLM-based system that closes this gap. It integrates vision, audio, tactile and force sensing into a reasoning–interaction–feedback loop: the model identifies missing information and chooses the target objects, interactions and modalities based on the instruction, while a physical interface executes it and returns the multi-sensory feedback.