Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

Oct 01, 2019

Ji Lin, Chuang Gan, Song Han

Figure 1 for Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

Figure 2 for Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

Figure 3 for Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

Figure 4 for Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

Share this with someone who'll enjoy it:

Abstract:Deep video recognition is more computationally expensive than image recognition, especially on large-scale datasets like Kinetics [1]. Therefore, training scalability is essential to handle a large amount of videos. In this paper, we study the factors that impact the training scalability of video networks. We recognize three bottlenecks, including data loading (data movement from disk to GPU), communication (data movement over networking), and computation FLOPs. We propose three design guidelines to improve the scalability: (1) fewer FLOPs and hardware-friendly operator to increase the computation efficiency; (2) fewer input frames to reduce the data movement and increase the data loading efficiency; (3) smaller model size to reduce the networking traffic and increase the networking efficiency. With these guidelines, we designed a new operator Temporal Shift Module (TSM) that is efficient and scalable for distributed training. TSM model can achieve 1.8x higher throughput compared to previous I3D models. We scale up the training of the TSM model to 1,536 GPUs, with a mini-batch of 12,288 video clips/98,304 images, without losing the accuracy. With such hardware-aware model design, we are able to scale up the training on Summit supercomputer and reduce the training time on Kinetics dataset from 49 hours 55 minutes to 14 minutes 13 seconds, achieving a top-1 accuracy of 74.0%, which is 1.6x and 2.9x faster than previous 3D video models with higher accuracy.

View paper on

Share this with someone who'll enjoy it:

Title:Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos

Paper and Code