Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Lei Wu

Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling

Jan 13, 2025

Jiebin Yan, Lei Wu, Yuming Fang, Xuelin Liu, Xue Xia, Weide Liu

Figure 1 for Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling

Figure 2 for Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling

Figure 3 for Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling

Figure 4 for Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling

Abstract:With the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from designing efficient video quality mapping models to various research directions, in-depth exploration of the effectiveness-efficiency trade-offs of spatio-temporal modeling in VQA models is still less sufficient. Considering the fact that videos have highly redundant information, this paper investigates this problem from the perspective of joint spatial and temporal sampling, aiming to seek the answer to how little information we should keep at least when feeding videos into the VQA models while with acceptable performance sacrifice. To this end, we drastically sample the video's information from both spatial and temporal dimensions, and the heavily squeezed video is then fed into a stable VQA model. Comprehensive experiments regarding joint spatial and temporal sampling are conducted on six public video quality databases, and the results demonstrate the acceptable performance of the VQA model when throwing away most of the video information. Furthermore, with the proposed joint spatial and temporal sampling strategy, we make an initial attempt to design an online VQA model, which is instantiated by as simple as possible a spatial feature extractor, a temporal feature fusion module, and a global quality regression module. Through quantitative and qualitative experiments, we verify the feasibility of online VQA model by simplifying itself and reducing input.

Via

Access Paper or Ask Questions

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Nov 16, 2024

Xudong Lu, Yinghao Chen, Cheng Chen, Hui Tan, Boheng Chen, Yina Xie, Rui Hu, Guanxin Tan, Renshou Wu, Yan Hu(+12 more)

Figure 1 for BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Figure 2 for BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Figure 3 for BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Figure 4 for BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Abstract:The emergence and growing popularity of multimodal large language models (MLLMs) have significant potential to enhance various aspects of daily life, from improving communication to facilitating learning and problem-solving. Mobile phones, as essential daily companions, represent the most effective and accessible deployment platform for MLLMs, enabling seamless integration into everyday tasks. However, deploying MLLMs on mobile phones presents challenges due to limitations in memory size and computational capability, making it difficult to achieve smooth and real-time processing without extensive optimization. In this paper, we present BlueLM-V-3B, an algorithm and system co-design approach specifically tailored for the efficient deployment of MLLMs on mobile platforms. To be specific, we redesign the dynamic resolution scheme adopted by mainstream MLLMs and implement system optimization for hardware-aware deployment to optimize model inference on mobile phones. BlueLM-V-3B boasts the following key highlights: (1) Small Size: BlueLM-V-3B features a language model with 2.7B parameters and a vision encoder with 400M parameters. (2) Fast Speed: BlueLM-V-3B achieves a generation speed of 24.4 token/s on the MediaTek Dimensity 9300 processor with 4-bit LLM weight quantization. (3) Strong Performance: BlueLM-V-3B has attained the highest average score of 66.1 on the OpenCompass benchmark among models with $\leq$ 4B parameters and surpassed a series of models with much larger parameter sizes (e.g., MiniCPM-V-2.6, InternVL2-8B).

* 21 pages

Via

Access Paper or Ask Questions

Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation

Oct 30, 2024

Ruiyu Xiao, Lei Wu, Yuhang Gou, Weinan Zhang, Ting Liu

Figure 1 for Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation

Figure 2 for Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation

Figure 3 for Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation

Figure 4 for Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation

Abstract:Argumentative essay generation (AEG) aims to generate complete texts on specific controversial topics or debates. Although current AEG methods can generate individual opinions, they often overlook the high-level connections between these opinions. This often leads to the generated results being mired in logical confusion, unable to proof their own arguments effectively. The generated essay may present evidence that contradicts the claims or they may fail to assemble the claims into logical flow. In this paper, we present a unified two-stage framework: Proof-Enhancement and Self-Annotation (PESA) for AEG with a focus on logical enhancement. Specifically, we first construct pseudo-labels for logical information,claims and grounds, using a large language model. We then propose a tree planning approach that introduces proof principles and ensures logical consistency. Extensive experimental results show that, benefiting from proof principle guidance, PESA generates argumentative essays with better logical validity and persuasiveness than strong baseline models.

* EMNLP 2024

Via

Access Paper or Ask Questions

How Transformers Implement Induction Heads: Approximation and Optimization Analysis

Oct 15, 2024

Mingze Wang, Ruoxi Yu, Weinan E, Lei Wu

Figure 1 for How Transformers Implement Induction Heads: Approximation and Optimization Analysis

Abstract:Transformers have demonstrated exceptional in-context learning capabilities, yet the theoretical understanding of the underlying mechanisms remain limited. A recent work (Elhage et al., 2021) identified a "rich" in-context mechanism known as induction head, contrasting with "lazy" $n$-gram models that overlook long-range dependencies. In this work, we provide both approximation and optimization analyses of how transformers implement induction heads. In the approximation analysis, we formalize both standard and generalized induction head mechanisms, and examine how transformers can efficiently implement them, with an emphasis on the distinct role of each transformer submodule. For the optimization analysis, we study the training dynamics on a synthetic mixed target, composed of a 4-gram and an in-context 2-gram component. This setting enables us to precisely characterize the entire training process and uncover an {\em abrupt transition} from lazy (4-gram) to rich (induction head) mechanisms as training progresses.

* 39 pages

Via

Access Paper or Ask Questions

DTactive: A Vision-Based Tactile Sensor with Active Surface

Oct 10, 2024

Jikai Xu, Lei Wu, Changyi Lin, Ding Zhao, Huazhe Xu

Abstract:The development of vision-based tactile sensors has significantly enhanced robots' perception and manipulation capabilities, especially for tasks requiring contact-rich interactions with objects. In this work, we present DTactive, a novel vision-based tactile sensor with active surfaces. DTactive inherits and modifies the tactile 3D shape reconstruction method of DTact while integrating a mechanical transmission mechanism that facilitates the mobility of its surface. Thanks to this design, the sensor is capable of simultaneously performing tactile perception and in-hand manipulation with surface movement. Leveraging the high-resolution tactile images from the sensor and the magnetic encoder data from the transmission mechanism, we propose a learning-based method to enable precise angular trajectory control during in-hand manipulation. In our experiments, we successfully achieved accurate rolling manipulation within the range of [ -180{\deg},180{\deg} ] on various objects, with the root mean square error between the desired and actual angular trajectories being less than 12{\deg} on nine trained objects and less than 19{\deg} on three novel objects. The results demonstrate the potential of DTactive for in-hand object manipulation in terms of effectiveness, robustness and precision.

* Submitted to ICRA 2025

Via

Access Paper or Ask Questions

Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition

Jul 17, 2024

Mohamad Amin Mohamadi, Zhiyuan Li, Lei Wu, Danica J. Sutherland

Figure 1 for Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition

Figure 2 for Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition

Figure 3 for Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition

Figure 4 for Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition

Abstract:We present a theoretical explanation of the ``grokking'' phenomenon, where a model generalizes long after overfitting,for the originally-studied problem of modular addition. First, we show that early in gradient descent, when the ``kernel regime'' approximately holds, no permutation-equivariant model can achieve small population error on modular addition unless it sees at least a constant fraction of all possible data points. Eventually, however, models escape the kernel regime. We show that two-layer quadratic networks that achieve zero training loss with bounded $\ell_{\infty}$ norm generalize well with substantially fewer training points, and further show such networks exist and can be found by gradient descent with small $\ell_{\infty}$ regularization. We further provide empirical evidence that these networks as well as simple Transformers, leave the kernel regime only after initially overfitting. Taken together, our results strongly support the case for grokking as a consequence of the transition from kernel-like behavior to limiting behavior of gradient descent on deep networks.

* Accepted by ICML 2024

Via

Access Paper or Ask Questions

Improving Generalization and Convergence by Enhancing Implicit Regularization

May 31, 2024

Mingze Wang, Haotian He, Jinbo Wang, Zilin Wang, Guanhua Huang, Feiyu Xiong, Zhiyu Li, Weinan E, Lei Wu

Figure 1 for Improving Generalization and Convergence by Enhancing Implicit Regularization

Figure 2 for Improving Generalization and Convergence by Enhancing Implicit Regularization

Figure 3 for Improving Generalization and Convergence by Enhancing Implicit Regularization

Figure 4 for Improving Generalization and Convergence by Enhancing Implicit Regularization

Abstract:In this work, we propose an Implicit Regularization Enhancement (IRE) framework to accelerate the discovery of flat solutions in deep learning, thereby improving generalization and convergence. Specifically, IRE decouples the dynamics of flat and sharp directions, which boosts the sharpness reduction along flat directions while maintaining the training stability in sharp directions. We show that IRE can be practically incorporated with {\em generic base optimizers} without introducing significant computational overload. Experiments show that IRE consistently improves the generalization performance for image classification tasks across a variety of benchmark datasets (CIFAR-10/100, ImageNet) and models (ResNets and ViTs). Surprisingly, IRE also achieves a $2\times$ {\em speed-up} compared to AdamW in the pre-training of Llama models (of sizes ranging from 60M to 229M) on datasets including Wikitext-103, Minipile, and Openwebtext. Moreover, we provide theoretical guarantees, showing that IRE can substantially accelerate the convergence towards flat minima in Sharpness-aware Minimization (SAM).

* 35 pages

Via

Access Paper or Ask Questions

Exploring Neural Network Landscapes: Star-Shaped and Geodesic Connectivity

Apr 09, 2024

Zhanran Lin, Puheng Li, Lei Wu

Figure 1 for Exploring Neural Network Landscapes: Star-Shaped and Geodesic Connectivity

Figure 2 for Exploring Neural Network Landscapes: Star-Shaped and Geodesic Connectivity

Figure 3 for Exploring Neural Network Landscapes: Star-Shaped and Geodesic Connectivity

Figure 4 for Exploring Neural Network Landscapes: Star-Shaped and Geodesic Connectivity

Abstract:One of the most intriguing findings in the structure of neural network landscape is the phenomenon of mode connectivity: For two typical global minima, there exists a path connecting them without barrier. This concept of mode connectivity has played a crucial role in understanding important phenomena in deep learning. In this paper, we conduct a fine-grained analysis of this connectivity phenomenon. First, we demonstrate that in the overparameterized case, the connecting path can be as simple as a two-piece linear path, and the path length can be nearly equal to the Euclidean distance. This finding suggests that the landscape should be nearly convex in a certain sense. Second, we uncover a surprising star-shaped connectivity: For a finite number of typical minima, there exists a center on minima manifold that connects all of them simultaneously via linear paths. These results are provably valid for linear networks and two-layer ReLU networks under a teacher-student setup, and are empirically supported by models trained on MNIST and CIFAR-10.

* The first two authors contributed equally

Via

Access Paper or Ask Questions

A Duality Analysis of Kernel Ridge Regression in the Noiseless Regime

Feb 24, 2024

Jihao Long, Xiaojun Peng, Lei Wu

Abstract:In this paper, we conduct a comprehensive analysis of generalization properties of Kernel Ridge Regression (KRR) in the noiseless regime, a scenario crucial to scientific computing, where data are often generated via computer simulations. We prove that KRR can attain the minimax optimal rate, which depends on both the eigenvalue decay of the associated kernel and the relative smoothness of target functions. Particularly, when the eigenvalue decays exponentially fast, KRR achieves the spectral accuracy, i.e., a convergence rate faster than any polynomial. Moreover, the numerical experiments well corroborate our theoretical findings. Our proof leverages a novel extension of the duality framework introduced by Chen et al. (2023), which could be useful in analyzing kernel-based methods beyond the scope of this work.

Via

Access Paper or Ask Questions

The Implicit Bias of Gradient Noise: A Symmetry Perspective

Feb 11, 2024

Liu Ziyin, Mingze Wang, Lei Wu

Figure 1 for The Implicit Bias of Gradient Noise: A Symmetry Perspective

Figure 2 for The Implicit Bias of Gradient Noise: A Symmetry Perspective

Figure 3 for The Implicit Bias of Gradient Noise: A Symmetry Perspective

Figure 4 for The Implicit Bias of Gradient Noise: A Symmetry Perspective

Abstract:We characterize the learning dynamics of stochastic gradient descent (SGD) when continuous symmetry exists in the loss function, where the divergence between SGD and gradient descent is dramatic. We show that depending on how the symmetry affects the learning dynamics, we can divide a family of symmetry into two classes. For one class of symmetry, SGD naturally converges to solutions that have a balanced and aligned gradient noise. For the other class of symmetry, SGD will almost always diverge. Then, we show that our result remains applicable and can help us understand the training dynamics even when the symmetry is not present in the loss function. Our main result is universal in the sense that it only depends on the existence of the symmetry and is independent of the details of the loss function. We demonstrate that the proposed theory offers an explanation of progressive sharpening and flattening and can be applied to common practical problems such as representation normalization, matrix factorization, and the use of warmup.

* preprint

Via

Access Paper or Ask Questions