Mobile manipulation in dynamic indoor scenes

DREAM: Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation

Zhijie Yan1 Shufei Li2 Ze Zhang1 Xin Liu1 Yuhang Zheng3 Zuoxu Wang1
1Beihang University 2City University of Hong Kong 3National University of Singapore

DREAM builds and updates spatio-semantic memory during task execution, allowing a mobile robot to search, navigate, grasp, and place objects in unfamiliar indoor environments.

DREAM framework executing a dynamic mobile manipulation task
DREAM actively acquires targets, reacquires relocated objects, and updates task-relevant scene memory during long-horizon manipulation.

Abstract

Reliable mobile manipulation in dynamic indoor environments requires a scene representation that remains geometrically consistent, semantically queryable, and computationally manageable. As objects move and camera poses are corrected, the robot must update its remembered scene while controlling the cost of retained observations and semantic features.

DREAM starts in an unfamiliar indoor environment without a pre-built map. It builds spatio-semantic voxel memory from RGB-D observations registered by a LiDAR-inertial-visual SLAM backend. Pose-graph-aware reintegration and Redundancy-Aware Memory Pruning maintain this memory, while language-conditioned retrieval, image detection, and semantic verification support target search and reacquisition. Navigation, grasping, and placement complete the task.

Highlights

01

No Pre-Built Map

DREAM operates in dynamic, previously unseen indoor environments from a natural-language instruction.

02

Dynamic Memory

New observations and corrected poses update the voxel memory; pruning limits retained observation history.

03

Hybrid Localization

Language-conditioned retrieval, image detection, and visual verification localize task objects.

04

Real Robot Tasks

The system is evaluated on navigation, pickup, placement, and long-horizon mobile manipulation.

Primary videos

Real-Robot Demonstrations

The two complete task recordings show robot execution, localization, and semantic memory together. Additional manipulation tasks and a separate SLAM recording are shown below.

Main demo 01

Scenario 01: Pick-and-place task

A complete manipulation sequence with synchronized views of the robot, localization, and memory.

Main demo 02

Scenario 02: Yellow knife to red basket

The target moves during approach. Fresh observations update its remembered location before manipulation continues.

More videos

Additional Dynamic Manipulation Results

Further examples of navigation, target search, and manipulation.

SLAM

SLAM Test in a 100 x 50 m Scene

A separate localization demonstration in a 100 × 50 m indoor scene, using sensor transforms obtained from the robot CAD model.

Additional views

Individual Camera and Memory Views

Separate recordings of robot execution, SLAM, and the memory interface for the two scenarios above.

01

Scenario 01 Breakdown

Three source views used to explain the first combination demo.

Third-person robot execution
SLAM and localization view
Server-side memory and task state
02

Scenario 02 Breakdown

Three source views used to explain the second combination demo.

Third-person robot execution
SLAM and localization view
Server-side memory and task state

Supplementary simulation

Cross-Room Dynamic Pick-and-Place

All 50 task recordings from the current residential evaluation: 38 completed tasks and 12 failures. Complete timelines play at 12×, with a current first-person view, saved observations, semantic memory, and navigation paths. The simulation page also reports the separate extended-search cases.

View all 50 simulation trials Simulation code and protocol

Method

Dynamic Memory, Hybrid Localization, and Task-Oriented Navigation

DREAM is organized around a closed loop across perception, memory, localization, navigation, and manipulation.

Hardware platform: AgileX Ranger Mini 3.0 mobile base, UFACTORY xArm6 manipulator, wrist-mounted Intel RealSense D435i RGB-D camera, and Livox MID-360 LiDAR with a built-in IMU.

Overview of the DREAM system architecture
System overview. DREAM builds and updates dynamic spatio-semantic memory while planning actions around task-relevant objects. The robot CAD assets are available in the repository docs directory, including ROBOT.jpg and ROBOT.zip exported from SolidWorks.
Hybrid localization pipeline in DREAM
Language-conditioned memory retrieval, image detection, and verification locate task objects and check whether remembered positions remain valid.
Task-oriented navigation in DREAM
Task-oriented navigation prioritizes uncertain or task-relevant regions and selects manipulation-aware docking poses.

Experimental Results

Real-robot tasks, residential simulation, and the controlled memory comparison.

Real-robot evaluation

DREAM completes 50 of 80 tasks, compared with 39 of 80 for DynaMem. The Wilson 95% intervals are 51.5–72.3% and 38.1–59.5%, respectively; the aggregate difference is not significant at 0.05 (Fisher p = 0.111).

Read the physical evaluation

Residential simulation

The recorded controller completes 38 of 50 tasks (76%), using seed 42, an 1800-second robot-action budget, and no fixed server deadline. All 38 successes pass independent physics, observation, and arm-return checks; all 12 failures are included. These houses were used during development, so this rate describes the evaluated cohort.

Simulation results and recordings

Dynamic and static memory

With perception, navigation, and Fetch manipulation held fixed, dynamic memory completes 22/30 tasks and static memory 14/30. The difference is 26.7 percentage points, with a paired house-level 95% bootstrap interval of 6.7–50.0 points.

View the comparison records

BibTeX

If you use this work in your research, please cite the paper.

BibTeX
@misc{yan2026dynamicresilientspatiosemanticmemory,
  title={Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation},
  author={Zhijie Yan and Shufei Li and Ze Zhang and Xin Liu and Yuhang Zheng and Zuoxu Wang},
  year={2026},
  eprint={2606.00576},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2606.00576}
}