Picture for Jian Xue

Jian Xue

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

Add code
Aug 04, 2026
Viaarxiv icon

IoUPD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

Add code
Jul 17, 2026
Viaarxiv icon

Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

Add code
Jul 14, 2026
Viaarxiv icon

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

Add code
Jun 29, 2026
Viaarxiv icon

QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval

Add code
Jun 18, 2026
Viaarxiv icon

Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression

Add code
Mar 04, 2026
Viaarxiv icon

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

Add code
Feb 10, 2026
Viaarxiv icon

Generative AI for Analysts

Add code
Dec 12, 2025
Viaarxiv icon

PHRASED: Phrase Dictionary Biasing for Speech Translation

Add code
Jun 10, 2025
Viaarxiv icon

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation

Add code
Feb 04, 2025
Figure 1 for Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Figure 2 for Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Figure 3 for Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Figure 4 for Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Viaarxiv icon