From 2.5D to 3D: Unprojecting RGB-D Sensor Data for Spatial Memory (Part 1)

In my last post evaluating SigLIP 2, the takeaway was pretty clear: pure 2D appearance models fail hard when you start rearranging a room. Global VLMs over-index on object semantics (like matching a red scoop) instead of understanding the actual 3D geometry of the place itself. If we want an agent to actually understand spatial relationships—or if we want to synthetically edit a scene—we need to lift our data out of the 2.5D image plane and into a true 3D coordinate space. We need continuous geometry. ...

August 20, 2026 · 3 min