Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

May 23, 2025

Ming Hu, Zhendi Yu, Feilong Tang, Kaiwen Chen, Yulong Li, Imran Razzak, Junjun He, Tolga Birdal, Kaijing Zhou, Zongyuan Ge

Figure 1 for Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

Figure 2 for Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

Figure 3 for Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

Figure 4 for Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

Share this with someone who'll enjoy it:

Abstract:Accurate 3D reconstruction of hands and instruments is critical for vision-based analysis of ophthalmic microsurgery, yet progress has been hampered by the lack of realistic, large-scale datasets and reliable annotation tools. In this work, we introduce OphNet-3D, the first extensive RGB-D dynamic 3D reconstruction dataset for ophthalmic surgery, comprising 41 sequences from 40 surgeons and totaling 7.1 million frames, with fine-grained annotations of 12 surgical phases, 10 instrument categories, dense MANO hand meshes, and full 6-DoF instrument poses. To scalably produce high-fidelity labels, we design a multi-stage automatic annotation pipeline that integrates multi-view data observation, data-driven motion prior with cross-view geometric consistency and biomechanical constraints, along with a combination of collision-aware interaction constraints for instrument interactions. Building upon OphNet-3D, we establish two challenging benchmarks-bimanual hand pose estimation and hand-instrument interaction reconstruction-and propose two dedicated architectures: H-Net for dual-hand mesh recovery and OH-Net for joint reconstruction of two-hand-two-instrument interactions. These models leverage a novel spatial reasoning module with weak-perspective camera modeling and collision-aware center-based representation. Both architectures outperform existing methods by substantial margins, achieving improvements of over 2mm in Mean Per Joint Position Error (MPJPE) and up to 23% in ADD-S metrics for hand and instrument reconstruction, respectively.

View paper on

Share this with someone who'll enjoy it:

Title:Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

Paper and Code