Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Jonathan Lorraine

Motion Attribution for Video Generation

Jan 13, 2026

Xindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba, Laura Leal-Taixé, Olga Russakovsky, Sanja Fidler, Jonathan Lorraine

Abstract:Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework that scales to modern, large, high-quality video datasets and models. We use this to study which fine-tuning clips improve or degrade temporal dynamics. Motive isolates temporal dynamics from static appearance via motion-weighted loss masks, yielding efficient and scalable motion-specific influence computation. On text-to-video models, Motive identifies clips that strongly affect motion and guides data curation that improves temporal consistency and physical plausibility. With Motive-selected high-influence data, our method improves both motion smoothness and dynamic degree on VBench, achieving a 74.1% human preference win rate compared with the pretrained base model. To our knowledge, this is the first framework to attribute motion rather than visual appearance in video generative models and to use it to curate fine-tuning data.

* See the project website at https://research.nvidia.com/labs/sil/projects/MOTIVE/

Via

Access Paper or Ask Questions

Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond

May 07, 2025

Jessie Richter-Powell, Antonio Torralba, Jonathan Lorraine

Abstract:We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of distilling a powerful generative prior into a separate parametric representation extends to the audio domain. Leveraging a single pretrained model, Audio-SDS enables a broad range of tasks without requiring specialized datasets. In particular, we demonstrate how Audio-SDS can guide physically informed impact sound simulations, calibrate FM-synthesis parameters, and perform prompt-specified source separation. Our findings illustrate the versatility of distillation-based methods across modalities and establish a robust foundation for future work using generative priors in audio tasks.

* See the project website at https://research.nvidia.com/labs/toronto-ai/Audio-SDS/

Via

Access Paper or Ask Questions

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Nov 14, 2024

Zhengyi Wang, Jonathan Lorraine, Yikai Wang, Hang Su, Jun Zhu, Sanja Fidler, Xiaohui Zeng

Figure 1 for LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Figure 2 for LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Figure 3 for LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Figure 4 for LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Abstract:This work explores expanding the capabilities of large language models (LLMs) pretrained on text to generate 3D meshes within a unified model. This offers key advantages of (1) leveraging spatial knowledge already embedded in LLMs, derived from textual sources like 3D tutorials, and (2) enabling conversational 3D generation and mesh understanding. A primary challenge is effectively tokenizing 3D mesh data into discrete tokens that LLMs can process seamlessly. To address this, we introduce LLaMA-Mesh, a novel approach that represents the vertex coordinates and face definitions of 3D meshes as plain text, allowing direct integration with LLMs without expanding the vocabulary. We construct a supervised fine-tuning (SFT) dataset enabling pretrained LLMs to (1) generate 3D meshes from text prompts, (2) produce interleaved text and 3D mesh outputs as required, and (3) understand and interpret 3D meshes. Our work is the first to demonstrate that LLMs can be fine-tuned to acquire complex spatial knowledge for 3D mesh generation in a text-based format, effectively unifying the 3D and text modalities. LLaMA-Mesh achieves mesh generation quality on par with models trained from scratch while maintaining strong text generation performance.

* See the project website at https://research.nvidia.com/labs/toronto-ai/LLaMA-Mesh/

Via

Access Paper or Ask Questions

Multi-student Diffusion Distillation for Better One-step Generators

Oct 30, 2024

Yanke Song, Jonathan Lorraine, Weili Nie, Karsten Kreis, James Lucas

Figure 1 for Multi-student Diffusion Distillation for Better One-step Generators

Figure 2 for Multi-student Diffusion Distillation for Better One-step Generators

Figure 3 for Multi-student Diffusion Distillation for Better One-step Generators

Figure 4 for Multi-student Diffusion Distillation for Better One-step Generators

Abstract:Diffusion models achieve high-quality sample generation at the cost of a lengthy multistep inference procedure. To overcome this, diffusion distillation techniques produce student generators capable of matching or surpassing the teacher in a single step. However, the student model's inference speed is limited by the size of the teacher architecture, preventing real-time generation for computationally heavy applications. In this work, we introduce Multi-Student Distillation (MSD), a framework to distill a conditional teacher diffusion model into multiple single-step generators. Each student generator is responsible for a subset of the conditioning data, thereby obtaining higher generation quality for the same capacity. MSD trains multiple distilled students, allowing smaller sizes and, therefore, faster inference. Also, MSD offers a lightweight quality boost over single-student distillation with the same architecture. We demonstrate MSD is effective by training multiple same-sized or smaller students on single-step distillation using distribution matching and adversarial distillation techniques. With smaller students, MSD gets competitive results with faster inference for single-step generation. Using 4 same-sized students, MSD sets a new state-of-the-art for one-step image generation: FID 1.20 on ImageNet-64x64 and 8.20 on zero-shot COCO2014.

* Project page: https://research.nvidia.com/labs/toronto-ai/MSD/

Via

Access Paper or Ask Questions

JacNet: Learning Functions with Structured Jacobians

Aug 23, 2024

Jonathan Lorraine, Safwan Hossain

Figure 1 for JacNet: Learning Functions with Structured Jacobians

Figure 2 for JacNet: Learning Functions with Structured Jacobians

Figure 3 for JacNet: Learning Functions with Structured Jacobians

Abstract:Neural networks are trained to learn an approximate mapping from an input domain to a target domain. Incorporating prior knowledge about true mappings is critical to learning a useful approximation. With current architectures, it is challenging to enforce structure on the derivatives of the input-output mapping. We propose to use a neural network to directly learn the Jacobian of the input-output function, which allows easy control of the derivative. We focus on structuring the derivative to allow invertibility and also demonstrate that other useful priors, such as $k$-Lipschitz, can be enforced. Using this approach, we can learn approximations to simple functions that are guaranteed to be invertible and easily compute the inverse. We also show similar results for 1-Lipschitz functions.

* 6 pages, 3 Figures, ICML 2019 INNF Workshop

Via

Access Paper or Ask Questions

Scalable Nested Optimization for Deep Learning

Jul 01, 2024

Jonathan Lorraine

Figure 1 for Scalable Nested Optimization for Deep Learning

Figure 2 for Scalable Nested Optimization for Deep Learning

Figure 3 for Scalable Nested Optimization for Deep Learning

Figure 4 for Scalable Nested Optimization for Deep Learning

Abstract:Gradient-based optimization has been critical to the success of machine learning, updating a single set of parameters to minimize a single loss. A growing number of applications rely on a generalization of this, where we have a bilevel or nested optimization of which subsets of parameters update on different objectives nested inside each other. We focus on motivating examples of hyperparameter optimization and generative adversarial networks. However, naively applying classical methods often fails when we look at solving these nested problems on a large scale. In this thesis, we build tools for nested optimization that scale to deep learning setups.

* View more research details at https://www.jonlorraine.com/

Via

Access Paper or Ask Questions

Improving Hyperparameter Optimization with Checkpointed Model Weights

Jun 26, 2024

Nikhil Mehta, Jonathan Lorraine, Steve Masson, Ramanathan Arunachalam, Zaid Pervaiz Bhat, James Lucas, Arun George Zachariah

Figure 1 for Improving Hyperparameter Optimization with Checkpointed Model Weights

Figure 2 for Improving Hyperparameter Optimization with Checkpointed Model Weights

Figure 3 for Improving Hyperparameter Optimization with Checkpointed Model Weights

Figure 4 for Improving Hyperparameter Optimization with Checkpointed Model Weights

Abstract:When training deep learning models, the performance depends largely on the selected hyperparameters. However, hyperparameter optimization (HPO) is often one of the most expensive parts of model design. Classical HPO methods treat this as a black-box optimization problem. However, gray-box HPO methods, which incorporate more information about the setup, have emerged as a promising direction for more efficient optimization. For example, using intermediate loss evaluations to terminate bad selections. In this work, we propose an HPO method for neural networks using logged checkpoints of the trained weights to guide future hyperparameter selections. Our method, Forecasting Model Search (FMS), embeds weights into a Gaussian process deep kernel surrogate model, using a permutation-invariant graph metanetwork to be data-efficient with the logged network weights. To facilitate reproducibility and further research, we open-source our code at https://github.com/NVlabs/forecasting-model-search.

* See the project website at https://research.nvidia.com/labs/toronto-ai/FMS/

Via

Access Paper or Ask Questions

Training Data Attribution via Approximate Unrolled Differentiation

May 21, 2024

Juhan Bae, Wu Lin, Jonathan Lorraine, Roger Grosse

Figure 1 for Training Data Attribution via Approximate Unrolled Differentiation

Figure 2 for Training Data Attribution via Approximate Unrolled Differentiation

Figure 3 for Training Data Attribution via Approximate Unrolled Differentiation

Figure 4 for Training Data Attribution via Approximate Unrolled Differentiation

Abstract:Many training data attribution (TDA) methods aim to estimate how a model's behavior would change if one or more data points were removed from the training set. Methods based on implicit differentiation, such as influence functions, can be made computationally efficient, but fail to account for underspecification, the implicit bias of the optimization algorithm, or multi-stage training pipelines. By contrast, methods based on unrolling address these issues but face scalability challenges. In this work, we connect the implicit-differentiation-based and unrolling-based approaches and combine their benefits by introducing Source, an approximate unrolling-based TDA method that is computed using an influence-function-like formula. While being computationally efficient compared to unrolling-based approaches, Source is suitable in cases where implicit-differentiation-based approaches struggle, such as in non-converged models and multi-stage training pipelines. Empirically, Source outperforms existing TDA techniques in counterfactual prediction, especially in settings where implicit-differentiation-based approaches fall short.

Via

Access Paper or Ask Questions

LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis

Mar 22, 2024

Kevin Xie, Jonathan Lorraine, Tianshi Cao, Jun Gao, James Lucas, Antonio Torralba, Sanja Fidler, Xiaohui Zeng

Figure 1 for LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis

Figure 2 for LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis

Figure 3 for LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis

Figure 4 for LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis

Abstract:Recent text-to-3D generation approaches produce impressive 3D results but require time-consuming optimization that can take up to an hour per prompt. Amortized methods like ATT3D optimize multiple prompts simultaneously to improve efficiency, enabling fast text-to-3D synthesis. However, they cannot capture high-frequency geometry and texture details and struggle to scale to large prompt sets, so they generalize poorly. We introduce LATTE3D, addressing these limitations to achieve fast, high-quality generation on a significantly larger prompt set. Key to our method is 1) building a scalable architecture and 2) leveraging 3D data during optimization through 3D-aware diffusion priors, shape regularization, and model initialization to achieve robustness to diverse and complex training prompts. LATTE3D amortizes both neural field and textured surface generation to produce highly detailed textured meshes in a single forward pass. LATTE3D generates 3D objects in 400ms, and can be further enhanced with fast test-time optimization.

* See the project website at https://research.nvidia.com/labs/toronto-ai/LATTE3D/

Via

Access Paper or Ask Questions

Graph Metanetworks for Processing Diverse Neural Architectures

Dec 07, 2023

Derek Lim, Haggai Maron, Marc T. Law, Jonathan Lorraine, James Lucas

Figure 1 for Graph Metanetworks for Processing Diverse Neural Architectures

Figure 2 for Graph Metanetworks for Processing Diverse Neural Architectures

Figure 3 for Graph Metanetworks for Processing Diverse Neural Architectures

Figure 4 for Graph Metanetworks for Processing Diverse Neural Architectures

Abstract:Neural networks efficiently encode learned information within their parameters. Consequently, many tasks can be unified by treating neural networks themselves as input data. When doing so, recent studies demonstrated the importance of accounting for the symmetries and geometry of parameter spaces. However, those works developed architectures tailored to specific networks such as MLPs and CNNs without normalization layers, and generalizing such architectures to other types of networks can be challenging. In this work, we overcome these challenges by building new metanetworks - neural networks that take weights from other neural networks as input. Put simply, we carefully build graphs representing the input neural networks and process the graphs using graph neural networks. Our approach, Graph Metanetworks (GMNs), generalizes to neural architectures where competing methods struggle, such as multi-head attention layers, normalization layers, convolutional layers, ResNet blocks, and group-equivariant linear layers. We prove that GMNs are expressive and equivariant to parameter permutation symmetries that leave the input neural network functions unchanged. We validate the effectiveness of our method on several metanetwork tasks over diverse neural network architectures.

* 29 pages

Via

Access Paper or Ask Questions