Picture for Steven Hoi

Steven Hoi

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

Add code
Aug 17, 2026
Viaarxiv icon

StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

Add code
Aug 03, 2026
Viaarxiv icon

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Add code
Jul 30, 2026
Viaarxiv icon

FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

Add code
Jul 23, 2026
Viaarxiv icon

Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA

Add code
Jul 13, 2026
Viaarxiv icon

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States

Add code
Jun 30, 2026
Viaarxiv icon

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

Add code
Jun 29, 2026
Viaarxiv icon

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

Add code
May 06, 2026
Viaarxiv icon

OMG-Avatar: One-shot Multi-LOD Gaussian Head Avatar

Add code
Mar 02, 2026
Viaarxiv icon

MAI-UI Technical Report: Real-World Centric Foundation GUI Agents

Add code
Dec 26, 2025
Figure 1 for MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
Figure 2 for MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
Figure 3 for MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
Figure 4 for MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
Viaarxiv icon