No Pre-Built Map
DREAM operates in dynamic, previously unseen indoor environments from a natural-language instruction.
Mobile manipulation in dynamic indoor scenes
DREAM builds and updates spatio-semantic memory during task execution, allowing a mobile robot to search, navigate, grasp, and place objects in unfamiliar indoor environments.
Reliable mobile manipulation in dynamic indoor environments requires a scene representation that remains geometrically consistent, semantically queryable, and computationally manageable. As objects move and camera poses are corrected, the robot must update its remembered scene while controlling the cost of retained observations and semantic features.
DREAM starts in an unfamiliar indoor environment without a pre-built map. It builds spatio-semantic voxel memory from RGB-D observations registered by a LiDAR-inertial-visual SLAM backend. Pose-graph-aware reintegration and Redundancy-Aware Memory Pruning maintain this memory, while language-conditioned retrieval, image detection, and semantic verification support target search and reacquisition. Navigation, grasping, and placement complete the task.
DREAM operates in dynamic, previously unseen indoor environments from a natural-language instruction.
New observations and corrected poses update the voxel memory; pruning limits retained observation history.
Language-conditioned retrieval, image detection, and visual verification localize task objects.
The system is evaluated on navigation, pickup, placement, and long-horizon mobile manipulation.
Primary videos
The two complete task recordings show robot execution, localization, and semantic memory together. Additional manipulation tasks and a separate SLAM recording are shown below.
Main demo 01
A complete manipulation sequence with synchronized views of the robot, localization, and memory.
Main demo 02
The target moves during approach. Fresh observations update its remembered location before manipulation continues.
More videos
Further examples of navigation, target search, and manipulation.
SLAM
A separate localization demonstration in a 100 × 50 m indoor scene, using sensor transforms obtained from the robot CAD model.
Additional views
Separate recordings of robot execution, SLAM, and the memory interface for the two scenarios above.
Three source views used to explain the first combination demo.
Three source views used to explain the second combination demo.
Supplementary simulation
All 50 task recordings from the current residential evaluation: 38 completed tasks and 12 failures. Complete timelines play at 12×, with a current first-person view, saved observations, semantic memory, and navigation paths. The simulation page also reports the separate extended-search cases.
Method
DREAM is organized around a closed loop across perception, memory, localization, navigation, and manipulation.
Hardware platform: AgileX Ranger Mini 3.0 mobile base, UFACTORY xArm6 manipulator, wrist-mounted Intel RealSense D435i RGB-D camera, and Livox MID-360 LiDAR with a built-in IMU.
ROBOT.jpg and
ROBOT.zip exported from SolidWorks.
Real-robot tasks, residential simulation, and the controlled memory comparison.
DREAM completes 50 of 80 tasks, compared with 39 of 80 for DynaMem. The Wilson 95% intervals are 51.5–72.3% and 38.1–59.5%, respectively; the aggregate difference is not significant at 0.05 (Fisher p = 0.111).
Read the physical evaluationThe recorded controller completes 38 of 50 tasks (76%), using seed 42, an 1800-second robot-action budget, and no fixed server deadline. All 38 successes pass independent physics, observation, and arm-return checks; all 12 failures are included. These houses were used during development, so this rate describes the evaluated cohort.
Simulation results and recordingsWith perception, navigation, and Fetch manipulation held fixed, dynamic memory completes 22/30 tasks and static memory 14/30. The difference is 26.7 percentage points, with a paired house-level 95% bootstrap interval of 6.7–50.0 points.
View the comparison recordsIf you use this work in your research, please cite the paper.
@misc{yan2026dynamicresilientspatiosemanticmemory,
title={Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation},
author={Zhijie Yan and Shufei Li and Ze Zhang and Xin Liu and Yuhang Zheng and Zuoxu Wang},
year={2026},
eprint={2606.00576},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2606.00576}
}