Picture for Takuma Yagi

Takuma Yagi

The Embodiment Gap in Robot Foundation Models

Add code
Aug 19, 2026
Viaarxiv icon

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

Add code
Aug 12, 2026
Viaarxiv icon

Learning Object States from Actions via Large Language Models

Add code
May 02, 2024
Figure 1 for Learning Object States from Actions via Large Language Models
Figure 2 for Learning Object States from Actions via Large Language Models
Figure 3 for Learning Object States from Actions via Large Language Models
Figure 4 for Learning Object States from Actions via Large Language Models
Viaarxiv icon

FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation

Add code
Feb 01, 2024
Figure 1 for FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
Figure 2 for FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
Figure 3 for FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
Figure 4 for FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
Viaarxiv icon

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Add code
Nov 30, 2023
Figure 1 for Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Figure 2 for Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Figure 3 for Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Figure 4 for Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Viaarxiv icon

Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos

Add code
Nov 29, 2023
Figure 1 for Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
Figure 2 for Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
Figure 3 for Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
Figure 4 for Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
Viaarxiv icon

Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos

Add code
Feb 10, 2023
Figure 1 for Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos
Figure 2 for Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos
Figure 3 for Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos
Figure 4 for Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos
Viaarxiv icon

Precise Affordance Annotation for Egocentric Action Video Datasets

Add code
Jun 11, 2022
Figure 1 for Precise Affordance Annotation for Egocentric Action Video Datasets
Figure 2 for Precise Affordance Annotation for Egocentric Action Video Datasets
Figure 3 for Precise Affordance Annotation for Egocentric Action Video Datasets
Figure 4 for Precise Affordance Annotation for Egocentric Action Video Datasets
Viaarxiv icon

Object Instance Identification in Dynamic Environments

Add code
Jun 10, 2022
Figure 1 for Object Instance Identification in Dynamic Environments
Figure 2 for Object Instance Identification in Dynamic Environments
Figure 3 for Object Instance Identification in Dynamic Environments
Figure 4 for Object Instance Identification in Dynamic Environments
Viaarxiv icon

Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction

Add code
Oct 19, 2021
Figure 1 for Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction
Figure 2 for Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction
Figure 3 for Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction
Figure 4 for Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction
Viaarxiv icon