Portrait of Peizhou Cao
Beijing ↔ Shanghai2026

About

Peizhou Cao.

Building models and worlds that help machines understand space.

I am a second-year Ph.D. student at Beihang University and Shanghai AI Laboratory, co-supervised by Jiangmiao Pang and Dahua Lin. My work focuses on spatial intelligence, multimodal foundation models, and world models for understanding and simulating visual environments.

Spatial intelligenceMultimodal foundation modelsWorld models & simulation

01 / Profile

Experience

Education

Ph.D. in Computer Science and Technology

Beihang University · Beijing

Research

Research Intern

Shanghai AI Laboratory · Shanghai

Education

B.Eng. in Electronic Engineering

Xidian University · Xi’an

2021 — 2022 National Scholarship

Spatial perceptionVision-language reasoningWorld modelsScalable simulation

02 / Research

Research landscape

Models, data, benchmarks, and integrated systems across my work.

Research taxonomy

01 core · 04 categories · 09 papers

Layer 01 · CoreSpatial intelligencePerceive → reason → simulate
Research architecture Spatial intelligence

Papers are organized by their primary role in the research pipeline.

Model Data Benchmark Integrated systems

03 / Selected work

Research, built to be used.

Nine selected works spanning models, benchmarks, datasets, and large-scale intelligent systems.

Hand-drawn PerceptionBench overview preserving the existing-benchmark example and ten atomic visual perception capabilities
Atomic visual perception
02arXiv · 2026

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Z. Lin, Y. Xie, B. Qu, H. Wang, J. Li, H. Wu, …, P. Cao, et al.

G²TAM teaser with six panels comparing static and dynamic scene tracking from text, point, and box prompts
Geometry-grounded perception
03 ICML · 2026

G²TAM: Geometry Grounded Track Anything Model

C. Zhu, P. Cao, J. Lin, W. Hu, Y. Ran, J. Pang, T. Wang, X. Liu

Paper ↗ Code coming soon
Hand-drawn Astra reasoning trajectory that observes two room views, rotates twice with a world simulator, and answers East
Reasoning through imagined worlds
04arXiv · 2026

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

C. Zhu*, J. Lin*, Y. Long, P. Cao, T. Wang, J. Pang, X. Liu

MMSI-Video-Bench teaser illustrating spatial construction, spatial reasoning, planning, and benchmark categories
Video spatial intelligence
05 arXiv · 2026

MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence

J. Lin*, R. Xu*, S. Zhu, S. Yang, P. Cao, et al.

SynthVerse teaser collage of animated, embodied, articulated, and manipulation sequences with colored point trajectories
Synthetic worlds in motion
06 SIGGRAPH · 2026

SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point Tracking

W. Zhao, H. Xu, X. Miao, Q. Zhao, R. Zhang, K. Huang, N. Gao, P. Cao, et al.

InternScenes teaser showing layout diversity, 1.9 million objects across 288 classes, layout generation, and embodied AI applications
Simulatable indoor worlds
07 NeurIPS D&B · 2025

InternScenes: A Large-Scale Simulatable Indoor Scene Dataset with Realistic Layouts

P. Cao*, W. Zhong*, Y. Jin, L. Luo, W. Cai, J. Lin, et al. (*Equal Contribution)

GRUtopia teaser showing more than 100,000 scenes across 89 categories and robot tasks in navigation, dialogue, and manipulation
Robots at city scale
08 arXiv · 2024

GRUtopia: Dream General Robots in a City at Scale

H. Wang, J. Chen, W. Huang, Q. Ben, T. Wang, B. Mi, …, P. Cao, et al.

09 IEEE T-CYB · 2026

A Progressive Semi-Distillation Model for Dual-Source Remote Sensing Image Classification

H. Zhu, P. Cao, L. Jiao, X. Li, B. Hou, X. Yi, W. Zhao, W. Ma

04 / Contact

Hello

Research conversations, open-source collaborations, and ambitious problems are always welcome.