Hi, I’m Sabarish Kuduwa.
I’m a Technical Lead on Document Understanding at Uber. I work on spatial models and Vision-Language Models that power large-scale document transcription systems, with a focus on getting production accuracy past human baselines.
Before that I was Technical Lead at the robotics startup eBots, where I built the core perception stack—from 3D reconstruction and pose estimation through the visual feedback loops used in embodied assembly systems.
My work sits at the intersection of multimodal reasoning and physical space: how to make models respect spatial constraints, and how to implement the lower-level C++/CUDA pieces that let machines actually perceive the world. This site is a clean-room place for the technical notes, architecture write-ups, and experiments that come out of that.
Core Interests
- Machine Learning & Deep Learning
- Computer Vision & 3D Reconstruction
- World Models & Embodied AI
- Robotics Perception
Feel free to reach out if you want to collaborate or discuss the math behind the metal