Picture for Jianlong Wu

Jianlong Wu

Key-Frame Reasoning with SAM3: Third Place Solution for the MeViS-Text Track of the 8th LSVOS Challenge

Add code
Aug 18, 2026
Viaarxiv icon

VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)

Add code
Aug 13, 2026
Viaarxiv icon

AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

Add code
Aug 13, 2026
Viaarxiv icon

Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning

Add code
Jun 14, 2026
Viaarxiv icon

Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge

Add code
Apr 01, 2026
Viaarxiv icon

The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation

Add code
Apr 01, 2026
Viaarxiv icon

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval

Add code
Jan 28, 2026
Viaarxiv icon

AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Add code
Mar 16, 2025
Viaarxiv icon

MegaSR: Mining Customized Semantics and Expressive Guidance for Image Super-Resolution

Add code
Mar 11, 2025
Viaarxiv icon

HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models

Add code
Feb 28, 2025
Figure 1 for HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
Figure 2 for HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
Figure 3 for HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
Figure 4 for HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
Viaarxiv icon