ROMA LLM System for Real-World Object-Centric Multi-Sensory Active Perception

I saw. I touched. I understood.

  1. 1Gaoling School of Artificial Intelligence, Renmin University of China
  2. 2Beijing Key Laboratory of Research on Large Models and Intelligent Governance
  3. 3Beijing Academy of Artificial Intelligence
  4. 4Beijing Jiaotong University
  5. 5Peking University
  6. 6AresoX

*Equal contribution  ·  ✉Corresponding author

Overview

From passive sensory integration to active perception through interaction

Humans understand the physical world through an active process. When sensory evidence is insufficient, we interact: we lift an object to feel its weight, or shake it to hear what is inside. However, existing multi-sensory robot systems mostly integrate the inputs they are given, rather than actively acquire the evidence they are missing.

ROMA is an LLM-based system that closes this gap. It integrates vision, audio, tactile and force sensing into a reasoning–interaction–feedback loop: the model identifies missing information and chooses the target objects, interactions and modalities based on the instruction, while a physical interface executes it and returns the multi-sensory feedback.

  • Whatinformation is missing?
  • Howto acquire it?
  • Whenis the evidence sufficient?
  • ~2,000everyday objects and content combinations
  • 6atomic interactions
  • 4sensory modalities
    (vision, audio, touch, force)
  • 2,100benchmark tasks
  • 72.9%ROMA-7B success
    (best baseline 53.0%)

Dataset

ROMI-2K: Real-World Multi-Sensory Object Interaction Dataset

Existing datasets trade object diversity against interaction diversity. ROMI-2K covers nearly 2,000 everyday objects and content combinations, each with six atomic interactions and synchronized vision, audio, tactile and force feedback, collected two complementary ways.

Handheld collection

1,657 object–content combinations

A modified multi-sensory UMI (Pika Sense with a microphone and two GelSight Mini sensors). Every object is handled at 3 grasp locations, and inside contents are varied to capture diverse feedback.

  • Wrist + third view
  • Audio
  • Touch × 2
Tabletop scenes

500 scenes, 1–5 objects each

An xArm 6 arm with a Robotiq 2F-140 gripper grasps each object and executes all six interactions. We collect 400 scenes for training and 100 scenes for testing, with the test set containing 56 unseen objects.

  • RGB + depth
  • Audio
  • Touch × 2
  • 3-axis force

Dataset statistics

Object materials

Objects with contents inside

Data examples Action (Modality) · audio plays on click

Every object is annotated (Gemini 3.5 Flash, then manually verified) with identity, material, hardness, roughness, texture, contents, and the sensory modality that can reveal them.

System

ROMA System: A Reasoning–Interaction–Feedback Loop in the Real World

ROMA System has two parts: the multi-sensory LLM ROMA-7B, and a physical interface that executes interactions and collects multi-sensory feedback. ROMA-7B determines whether the current evidence is sufficient for the task. When additional evidence is required, it localizes the target object, selects the interaction and sensory modalities to acquire, and directs the physical interface to execute them. The collected feedback is then returned to ROMA-7B for further reasoning, forming an iterative perception-action loop until the task is completed.

ROMA system overview: grasp interface, six real-world interactions, multi-sensor feedback and the Multi-Sensory LLM ROMA-7B
Given an initial scene and an instruction, ROMA-7B identifies the target object and selects the interaction and sensory modalities to acquire. The physical interface then executes the selected actions and returns visual, audio, tactile, and force feedback.
1

Reasoning

ROMA-7B, built on Qwen2.5-Omni, learns to judge sufficiency and choose what to sense next.

Tokens

Action & modality tokens

Six action tokens, <grasp_start>/<grasp_end> around the target object, and positional tokens for touch and force. Each is initialized from the averaged embeddings of its words with Gaussian noise.

Stage 1

Multi-sensory alignment

Audio and touch are aligned to the vision–text space with audio-text, audio-visual-text, tactile-text and tactile-visual-text instruction tuning, so ambiguous signals are grounded in object semantics.

Stage 2

Active perception SFT

The model analyzes the scene, decides whether evidence is sufficient, and otherwise emits the next target object, interaction and modalities. It then reasons over the returned feedback and repeats.

Dynamic scene sampling

A smooth shift from touching things to solving scenes

Handheld and pseudo-scene samples draw positions from Beta(1, α), real-scene samples from Beta(α, 1), with α = 2.5. Sorting by position front-loads interaction-level supervision and leaves realistic scenes, and the force modality, for late training.

2

Interaction

The physical interface consists of two components: a robotic grasping module that securely holds the object, and a set of six predefined interactions that actively probe its physical properties.

Grasp interface

Stable contact for interaction

AnyGrasp proposes grasps from the object’s point cloud. A symmetry-based completion recovers geometry that transparent and reflective objects lose, then candidates are filtered by height, reachability and orientation, and refined for our gripper.

Object point cloud with completed points and the object, workspace and search cuboids
Point-cloud completion and grasp enhancement: completed points (purple) and the object, workspace and search cuboids around an object.
Grasp poses predicted by AnyGrasp (left) and by ROMA (right) on the same object
AnyGrasp pose vs. ROMA pose. The original AnyGrasp prediction (left) is sometimes an unstable top-down grasp that fails under vigorous interactions such as shaking and rotating. With point-cloud completion (purple) and grasp enhancement, ROMA (right) produces a stable grasp for interactions.

Six atomic interactions

  • <lift>Hold the object and raise it
  • <press>Press it down against the table
  • <collide>Tap it on the table repeatedly
  • <shake>Shake it from side to side
  • <rotate>Rotate it 90° about the grasp axis and rotate it back
  • <squeeze>Close the gripper a little further
3

Feedback

Every interaction is recorded by four sensory modalities; the model chooses which of them to read.

Wrist-camera image during an interaction

Vision

One wrist-camera image per interaction, showing the object’s appearance and how it deforms or moves.

Audio spectrogram of an interaction

Audio

One cropped microphone segment of the sound that the interaction makes.

GelSight tactile image

Touch

Two GelSight Mini images, one at the start and one during the interaction, showing local hardness, roughness and texture at the contact.

Force versus time during shaking

Force

A wrist force sensor. The mean gravity component during the interaction is given to the LLM as text, a cue for weight.

Force is moved to the robot base frame and the empty-gripper force is removed: Ftrans = RB←G RG←S (Fraw − F0raw)

Benchmark

ROMA Bench: Active Perception as Perception Chains

Active perception is formulated as a chain, where each interaction provides new evidence that informs subsequent decisions. ROMA Bench contains 2,100 scene-level tasks from the tabletop test split, organized into three types: single-chain, multi-chain, and intent-driven perception.

Single-chain

“Which objects are empty?”

Empty? shake A Answer

One target attribute, one chain of interactions.

Multi-chain

“Is the heaviest object also empty?”

Heavy? lift F Empty? shake A F reuse Answer

Several attributes. Chains share feedback, so redundant interactions should be skipped.

Intent-driven

“Find me lots of treats to eat.”

Instruction Heavy? lift F Inside? shake A F reuse Answer

The model infers which attributes matter from an implicit instruction, then chains them.

Tasks by type

Question formats

  • Multiple choice
  • True or false
  • Ranking
  • Unseen objects100 tabletop scenes, 2,100 tasks with 56 unseen objects, disjoint from the training data
  • Intent-driven questionswritten by GPT-5.4 from the scene annotations
  • Grasp success criterionthe predicted bbox must overlap a ground-truth by at least 70%
  • Real-robot execution132 free-form tasks, answer options removed

Which attributes appear together

Arc length: tasks involving the attribute. Ribbons: the two attributes linked by a Multi-chain task (1,036 tasks). The lighter part of each arc is Single-chain tasks (741), which involve one attribute.

Results

ROMA-7B beats frontier MLLMs on multi-sensory active perception

ROMA-7B demonstrates the strongest advantages on long-chain and multi-attribute tasks. Numbers are task success rates (%).

Success rate by task type

    Per-attribute success on ROMA Bench

    Har. hardness · Rou. roughness · Tex. texture · Ins. inside contents · Mat. material · Wei. weight. Darker is higher; bold marks the best in each column.

    +24.2Multi-chain
    over the best baseline
    +13.7Intent-driven
    over the best baseline
    +19.7Real robot
    over the best baseline

    Real-robot demos

    One real-world run for each task type, autonomously‌ executed by ROMA System on the robot arm. The videos include audio, so unmute them to hear the interaction sounds. Some waiting and moving segments are sped up.

    Single-chain

    TaskWhich object is the heaviest?

    0:00 / 0:00
    Multi-chain

    TaskIs the heaviest object empty?

    0:00 / 0:00
    Intent-driven

    TaskI am hungry.

    0:00 / 0:00

    Analysis

    How ROMA interacts, and why each sense matters

    Accuracy vs. interaction cost

    Interactions per scene on ROMA Bench with oracle boxes. Exhaustive grasps every object and runs all six interactions, about 4× the cost of ROMA-7B. GPT-5.4 and Gemini stop earlier, but often before the evidence is sufficient.

    Training ablations

    Drop in overall success (72.9%) when one component is removed.

    Share of each interaction

      Frontier models lean on <lift> and <squeeze>. ROMA uses <rotate>, <press> and <collide> as complementary probes.

      Remove a modality, the matching capability degrades

      Single-chain success rates when the modality returns “Not Available”. Overall: 72.9 → 64.1 without audio, 62.7 without touch, 62.2 without force.

      Case study

      Initial scene: a plush pig, a metal tin, and three objects on the right: a white box, a grey cube and a black box white boxgrey cubeblack box
      Initial scene

      QuestionWhich of the three rightmost objects are empty?

      GTthe white box

      GPT-5.4

      ✕ the grey cube

      <grasp> white box on the left <\grasp> (use touch and force) <shake>. The grey block is not a container. <grasp> black box on the right <\grasp> (use touch and force) <squeeze>. [the grey cube]

      • Inappropriate action–modality pairs
      • Lack of active interaction due to overconfidence

      Gemini 3.5 Flash

      ✕ the grey cube

      I need to interact with each of them (the white box, the grey cube, and the black box) by shaking them to hear if there is anything inside. <grasp> <\grasp> (use audio and force) <shake>. <grasp> <\grasp> (use audio and force) <shake>. <grasp> <\grasp> (use audio and force) <shake>. The answer is [the grey cube].

      • Inappropriate action–modality pairs

      ROMA-7B

      ✓ the white box

      <grasp> the white box <\grasp> (use audio) <shake> (use audio) <rotate>. … little noise. (use vision, touch and force) <squeeze>. The mean gravity force … similar to an empty version. <grasp> the grey block <\grasp> (use audio) <shake> The object produces … when shaken. <grasp> the black box <\grasp> (use audio) <shake>. The object produces … when shaken. The answer is the white box.

      • Continues interacting when evidence is insufficient
      • Selects complementary actions and modalities

      GPT-5.4 tends to stop with fewer interactions, occasionally skipping critical objects and sensory evidence. GPT-5.4 and Gemini 3.5 Flash also pair actions and modalities badly, for example trying to read gravity force during vigorous shaking.

      Citation

      BibTeX

      @article{feng2026roma,
        title   = {ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception},
        author  = {Feng, Ruoxuan and Chen, Yutong and Song, Ruihua and Yang, Huan and
                   Wang, Zhongyuan and Yao, Guocai and Hu, Di},
        journal = {arXiv preprint arXiv:2610.06955},
        url     = {https://arxiv.org/abs/2610.06955},
        year    = {2026}
      }