Picture for Jan Kautz

Jan Kautz

NVIDIA

MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Add code
Jul 10, 2024
Viaarxiv icon

An Empirical Study of Mamba-based Language Models

Add code
Jun 12, 2024
Viaarxiv icon

Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation

Add code
Jun 11, 2024
Figure 1 for Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
Figure 2 for Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
Figure 3 for Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
Figure 4 for Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
Viaarxiv icon

Flextron: Many-in-One Flexible Large Language Model

Add code
Jun 11, 2024
Viaarxiv icon

CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation

Add code
Jun 04, 2024
Figure 1 for CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Figure 2 for CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Figure 3 for CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Figure 4 for CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Viaarxiv icon

SpatialRGPT: Grounded Spatial Reasoning in Vision Language Model

Add code
Jun 03, 2024
Viaarxiv icon

X-VILA: Cross-Modality Alignment for Large Language Model

Add code
May 29, 2024
Viaarxiv icon

OmniDrive: A Holistic LLM-Agent Framework for Autonomous Driving with 3D Perception, Reasoning and Planning

Add code
May 02, 2024
Figure 1 for OmniDrive: A Holistic LLM-Agent Framework for Autonomous Driving with 3D Perception, Reasoning and Planning
Figure 2 for OmniDrive: A Holistic LLM-Agent Framework for Autonomous Driving with 3D Perception, Reasoning and Planning
Figure 3 for OmniDrive: A Holistic LLM-Agent Framework for Autonomous Driving with 3D Perception, Reasoning and Planning
Figure 4 for OmniDrive: A Holistic LLM-Agent Framework for Autonomous Driving with 3D Perception, Reasoning and Planning
Viaarxiv icon

LITA: Language Instructed Temporal-Localization Assistant

Add code
Mar 27, 2024
Figure 1 for LITA: Language Instructed Temporal-Localization Assistant
Figure 2 for LITA: Language Instructed Temporal-Localization Assistant
Figure 3 for LITA: Language Instructed Temporal-Localization Assistant
Figure 4 for LITA: Language Instructed Temporal-Localization Assistant
Viaarxiv icon

FoVA-Depth: Field-of-View Agnostic Depth Estimation for Cross-Dataset Generalization

Add code
Jan 24, 2024
Viaarxiv icon