I'm building intelligence at Amazon AGI, working on improving the reasoning ability of MLLMs and agentic systems.
My goal is to help models see and interact with the world, while making them useful for everyday interactions and scientific discovery.
Layer-wise study of FM priors for lighting robustness, geometric consistency, scene dynamics, and semantic awareness — measured via relighting PSNR/LPIPS, pose/flow error, and 3D semantic IoU.
Streaming attribution framework that tracks per-sample learning contributions to mitigate catastrophic forgetting, with real-time valuation for efficient replay across multi-task retention.
SFT and PPO alignment pipelines tuning LLMs for automated optical design and property prediction, validated against physical simulators on custom reasoning benchmarks.
Probing audiovisual integration in SOTA VLMs (GPT-4o, Claude 3.5, Gemini 1.5 Pro) across 2,556 questions on 700 videos. Open toolkit and public leaderboard for reproducible multimodal evaluation.
A lens-agnostic frontend that lifts pixels from any calibrated camera to the sphere for distortion-free, rotation-equivariant perception.
Paper →Diffusion model conditioned on geometric priors for joint estimation of camera distortions and scene geometry.
Paper →A family of foundation models for enterprise-grade reasoning, multimodal processing, and real-time conversational AI.
Report →2,556 audiovisual questions across 700 videos for evaluating tightly integrated audio-video reasoning in foundation models.
Paper →Dense optical flow via differentiable Gaussian rasterization with Lucas-Kanade alignment.
Paper →Neural refractive fields for modeling optical aberrations from protective covers.
Paper →Structured graph representations for understanding physical human-object interactions.
Paper →Graph attention networks guided by global shape priors for robust point cloud completion.
Paper →Hierarchical graph neural networks for efficient segmentation of large-scale 3D scenes.
Paper →