Picture for Benjia Zhou

Benjia Zhou

Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution

Add code
Aug 03, 2026
Viaarxiv icon

Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction

Add code
Jul 15, 2026
Viaarxiv icon

RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation

Add code
Dec 08, 2025
Viaarxiv icon

C${^2}$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval

Add code
Aug 19, 2024
Figure 1 for C${^2}$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval
Figure 2 for C${^2}$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval
Figure 3 for C${^2}$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval
Figure 4 for C${^2}$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval
Viaarxiv icon

Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation

Add code
Mar 19, 2024
Figure 1 for Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation
Figure 2 for Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation
Figure 3 for Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation
Figure 4 for Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation
Viaarxiv icon

PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features

Add code
Dec 05, 2023
Figure 1 for PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features
Figure 2 for PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features
Figure 3 for PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features
Figure 4 for PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features
Viaarxiv icon

Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition

Add code
Sep 11, 2023
Figure 1 for Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition
Figure 2 for Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition
Figure 3 for Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition
Figure 4 for Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition
Viaarxiv icon

Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining

Add code
Jul 27, 2023
Figure 1 for Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining
Figure 2 for Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining
Figure 3 for Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining
Figure 4 for Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining
Viaarxiv icon

A Unified Multimodal De- and Re-coupling Framework for RGB-D Motion Recognition

Add code
Nov 16, 2022
Viaarxiv icon

Effective Vision Transformer Training: A Data-Centric Perspective

Add code
Sep 29, 2022
Figure 1 for Effective Vision Transformer Training: A Data-Centric Perspective
Figure 2 for Effective Vision Transformer Training: A Data-Centric Perspective
Figure 3 for Effective Vision Transformer Training: A Data-Centric Perspective
Figure 4 for Effective Vision Transformer Training: A Data-Centric Perspective
Viaarxiv icon