Picture for Zhipeng Bao

Zhipeng Bao

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

Add code
Jul 27, 2026
Viaarxiv icon

DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception

Add code
Jun 24, 2026
Viaarxiv icon

CABLE: Cloud-Assisted Bandwidth-efficient LMM-based Encoding for V2X Systems

Add code
Jun 17, 2026
Viaarxiv icon

BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models

Add code
May 27, 2026
Viaarxiv icon

Walk through Paintings: Egocentric World Models from Internet Priors

Add code
Jan 21, 2026
Viaarxiv icon

Large Language Model-assisted Autonomous Vehicle Recovery from Immobilization

Add code
Oct 29, 2025
Viaarxiv icon

Your Ride, Your Rules: Psychology and Cognition Enabled Automated Driving Systems

Add code
Jun 13, 2025
Viaarxiv icon

Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models

Add code
Nov 07, 2024
Figure 1 for Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
Figure 2 for Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
Figure 3 for Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
Figure 4 for Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
Viaarxiv icon

ReferEverything: Towards Segmenting Everything We Can Speak of in Videos

Add code
Oct 30, 2024
Figure 1 for ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
Figure 2 for ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
Figure 3 for ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
Figure 4 for ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
Viaarxiv icon

Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding

Add code
Sep 05, 2024
Figure 1 for Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
Figure 2 for Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
Figure 3 for Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
Figure 4 for Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
Viaarxiv icon