Picture for Yifan Du

Yifan Du

Rotational Symmetry based Object Pose Estimation from Point Clouds in the Absence of Known 3D Models

Add code
Jun 15, 2026
Viaarxiv icon

Improving Vision-language Models with Perception-centric Process Reward Models

Add code
Apr 27, 2026
Viaarxiv icon

Towards Long-horizon Agentic Multimodal Search

Add code
Apr 14, 2026
Viaarxiv icon

VIPER: Process-aware Evaluation for Generative Video Reasoning

Add code
Dec 31, 2025
Viaarxiv icon

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

Add code
Jul 02, 2025
Viaarxiv icon

GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Add code
Jul 02, 2025
Figure 1 for GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Figure 2 for GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Figure 3 for GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Figure 4 for GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Viaarxiv icon

Seed1.5-VL Technical Report

Add code
May 11, 2025
Figure 1 for Seed1.5-VL Technical Report
Figure 2 for Seed1.5-VL Technical Report
Figure 3 for Seed1.5-VL Technical Report
Figure 4 for Seed1.5-VL Technical Report
Viaarxiv icon

Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Add code
Jan 03, 2025
Viaarxiv icon

Exploring the Design Space of Visual Context Representation in Video MLLMs

Add code
Oct 17, 2024
Figure 1 for Exploring the Design Space of Visual Context Representation in Video MLLMs
Figure 2 for Exploring the Design Space of Visual Context Representation in Video MLLMs
Figure 3 for Exploring the Design Space of Visual Context Representation in Video MLLMs
Figure 4 for Exploring the Design Space of Visual Context Representation in Video MLLMs
Viaarxiv icon

Towards Event-oriented Long Video Understanding

Add code
Jun 20, 2024
Figure 1 for Towards Event-oriented Long Video Understanding
Figure 2 for Towards Event-oriented Long Video Understanding
Figure 3 for Towards Event-oriented Long Video Understanding
Figure 4 for Towards Event-oriented Long Video Understanding
Viaarxiv icon