Picture for Yuanchen Wu

Yuanchen Wu

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

Add code
Jul 09, 2026
Viaarxiv icon

Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

Add code
Mar 21, 2026
Viaarxiv icon

Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration

Add code
Sep 17, 2025
Viaarxiv icon

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs

Add code
Jul 30, 2025
Viaarxiv icon

AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models

Add code
Jul 03, 2025
Figure 1 for AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
Figure 2 for AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
Figure 3 for AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
Figure 4 for AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
Viaarxiv icon

Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception

Add code
Apr 29, 2025
Figure 1 for Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
Figure 2 for Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
Figure 3 for Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
Figure 4 for Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
Viaarxiv icon

ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts

Add code
Apr 01, 2025
Figure 1 for ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts
Figure 2 for ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts
Figure 3 for ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts
Figure 4 for ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts
Viaarxiv icon

Robust Offline Active Learning on Graphs

Add code
Aug 15, 2024
Viaarxiv icon

DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation

Add code
Mar 17, 2024
Figure 1 for DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation
Figure 2 for DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation
Figure 3 for DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation
Figure 4 for DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation
Viaarxiv icon