Picture for Yonghui Wang

Yonghui Wang

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering

Add code
Jul 31, 2026
Viaarxiv icon

Hierarchical Evidence-Driven Reasoning for Long Document Understanding

Add code
Jul 06, 2026
Viaarxiv icon

Revisiting Shadow Detection from a Vision-Language Perspective

Add code
May 12, 2026
Viaarxiv icon

DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding

Add code
Aug 10, 2025
Viaarxiv icon

ROOT: VLM based System for Indoor Scene Understanding and Beyond

Add code
Nov 24, 2024
Figure 1 for ROOT: VLM based System for Indoor Scene Understanding and Beyond
Figure 2 for ROOT: VLM based System for Indoor Scene Understanding and Beyond
Figure 3 for ROOT: VLM based System for Indoor Scene Understanding and Beyond
Figure 4 for ROOT: VLM based System for Indoor Scene Understanding and Beyond
Viaarxiv icon

AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding

Add code
Aug 30, 2024
Figure 1 for AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
Figure 2 for AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
Figure 3 for AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
Figure 4 for AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
Viaarxiv icon

LaneTCA: Enhancing Video Lane Detection with Temporal Context Aggregation

Add code
Aug 25, 2024
Viaarxiv icon

SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection

Add code
Aug 07, 2024
Figure 1 for SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
Figure 2 for SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
Figure 3 for SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
Figure 4 for SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
Viaarxiv icon

TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Add code
Apr 15, 2024
Viaarxiv icon

Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs

Add code
Nov 22, 2023
Viaarxiv icon