Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Xiangru Lian

Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization

Jun 10, 2017
Xiangru Lian, Yijun Huang, Yuncheng Li, Ji Liu

Figure 1 for Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization

Figure 2 for Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization

Figure 3 for Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization

Figure 4 for Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization

Asynchronous parallel implementations of stochastic gradient (SG) have been broadly used in solving deep neural network and received many successes in practice recently. However, existing theories cannot explain their convergence and speedup properties, mainly due to the nonconvexity of most deep learning formulations and the asynchronous parallel mechanism. To fill the gaps in theory and provide theoretical supports, this paper studies two asynchronous parallel implementations of SG: one is on the computer network and the other is on the shared memory system. We establish an ergodic convergence rate $O(1/\sqrt{K})$ for both algorithms and prove that the linear speedup is achievable if the number of workers is bounded by $\sqrt{K}$ ($K$ is the total number of iterations). Our results generalize and improve existing analysis for convex minimization.

* 31 pages

Via

Access Paper or Ask Questions

Staleness-aware Async-SGD for Distributed Deep Learning

Apr 05, 2016
Wei Zhang, Suyog Gupta, Xiangru Lian, Ji Liu

Figure 1 for Staleness-aware Async-SGD for Distributed Deep Learning

Figure 2 for Staleness-aware Async-SGD for Distributed Deep Learning

Figure 3 for Staleness-aware Async-SGD for Distributed Deep Learning

Figure 4 for Staleness-aware Async-SGD for Distributed Deep Learning

Deep neural networks have been shown to achieve state-of-the-art performance in several machine learning tasks. Stochastic Gradient Descent (SGD) is the preferred optimization algorithm for training these networks and asynchronous SGD (ASGD) has been widely adopted for accelerating the training of large-scale deep networks in a distributed computing environment. However, in practice it is quite challenging to tune the training hyperparameters (such as learning rate) when using ASGD so as achieve convergence and linear speedup, since the stability of the optimization algorithm is strongly influenced by the asynchronous nature of parameter updates. In this paper, we propose a variant of the ASGD algorithm in which the learning rate is modulated according to the gradient staleness and provide theoretical guarantees for convergence of this algorithm. Experimental verification is performed on commonly-used image classification benchmarks: CIFAR10 and Imagenet to demonstrate the superior effectiveness of the proposed approach, compared to SSGD (Synchronous SGD) and the conventional ASGD algorithm.

* Accepted by IJCAI 2016

Via

Access Paper or Ask Questions