Building intelligence at Amazon AGI

Louise Xie

I'm building intelligence at Amazon AGI, working on improving the reasoning ability of MLLMs and agentic systems.
My goal is to help models see and interact with the world, while making them useful for everyday interactions and scientific discovery.

Louise Xie

Background

Applied Scientist · Amazon AGI
Post-training RL for Nova, video reasoning benchmarks, diffusion-based multimodal generation.
Student Researcher and PhD Fellowship · Google XR
Calibration foundation model for in-the-wild camera calibration. Efficient online stereo.
Ph.D. · Carnegie Mellon University
Robotics Institute, CUBE Lab. Computer vision, 3D reconstruction, neural representations.
BASc · University of Toronto
Mechanical Engineering, Robotics & Bioengineering minors. Manning Fellowship, NSERC USRA.

News

2026 USF accepted at CVPR 2026
2026 BEAM 2nd workshop accepted at ECCV 2026
2026 MAVERIX accepted to AAAI 2026
2025 AlignDiff accepted at ICCV 2025
2025 Contributed to Amazon Nova 2 multimodal reasoning & generation release
2025.06 Co-chaired BEAM 2025 workshop at CVPR 2025
2025.01 GSLK accepted at ICLR 2025
2024.10 SynthCover accepted to WACV 2025

Blog

Thoughts 2025.06.08

Hello World — Why I'm Starting a Blog

Notes on the self-evolution gaps in today's foundation models, and on writing as a way to think them through.

Current Research

In Submission Apr 2025 — Dec 2025

Probing Foundation Model Embeddings with 3D Gaussian Splatting

Layer-wise study of FM priors for lighting robustness, geometric consistency, scene dynamics, and semantic awareness — measured via relighting PSNR/LPIPS, pose/flow error, and 3D semantic IoU.

Ongoing Apr 2026 — Present

Continual Learning via Online Data Attribution

Streaming attribution framework that tracks per-sample learning contributions to mitigate catastrophic forgetting, with real-time valuation for efficient replay across multi-task retention.

Ongoing Jun 2025 — Present

LLM Post-training for Optical Discovery

SFT and PPO alignment pipelines tuning LLMs for automated optical design and property prediction, validated against physical simulators on custom reasoning benchmarks.

Recently Wrapped Sep 2024 — Mar 2025

MAVERIX — Multimodal Audio-Visual Evaluation Benchmark

Probing audiovisual integration in SOTA VLMs (GPT-4o, Claude 3.5, Gemini 1.5 Pro) across 2,556 questions on 700 videos. Open toolkit and public leaderboard for reproducible multimodal evaluation.

Selected Publications

CVPR 2026 2026

Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera

A lens-agnostic frontend that lifts pixels from any calibrated camera to the sphere for distortion-free, rotation-equivariant perception.

Paper →
ICCV 2025 2025

AlignDiff: Learning Physically-Grounded Camera Alignment via Diffusion

Diffusion model conditioned on geometric priors for joint estimation of camera distortions and scene geometry.

Paper →
Tech Report 2025

Amazon Nova 2: Multimodal Reasoning and Generation Models

A family of foundation models for enterprise-grade reasoning, multimodal processing, and real-time conversational AI.

Report →
AAAI 2026 2026

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

2,556 audiovisual questions across 700 videos for evaluating tightly integrated audio-video reasoning in foundation models.

Paper →
ICLR 2025 Jan 2025

Gaussian Splatting Lucas-Kanade

Dense optical flow via differentiable Gaussian rasterization with Lucas-Kanade alignment.

Paper →
WACV 2025 Oct 2024

Through the Curved Cover: Synthesizing Cover-Aberrated Scenes with Refractive Field

Neural refractive fields for modeling optical aberrations from protective covers.

Paper →
Computers in Industry 2025 2025

Learning 3D Human-Object Interaction Graphs

Structured graph representations for understanding physical human-object interactions.

Paper →
ACM SAC 2022

SAGA-Net: Shape-Assisted Graph Attention for Point-cloud Completion

Graph attention networks guided by global shape priors for robust point cloud completion.

Paper →
CoRL 2020 2020

MuGNet: Multi-Resolution GNN for Large-Scale Point-cloud Segmentation

Hierarchical graph neural networks for efficient segmentation of large-scale 3D scenes.

Paper →