Seminar第2188讲 残差网络的多卡并行训练方法

创建时间:  2021/11/16  谭福平   浏览次数:   返回

报告题目 (Title):Layer-Parallel Training of Residual Networks with Auxiliary-Variable Networks (残差网络的多卡并行训练方法)

报告人 (Speaker): 孙琪 助理教授(同济大学)

报告时间 (Time):2021年11月18日 (周四) 10:55-11:40

报告地点 (Place):宝山校区D123

邀请人(Inviter):朱翔、潘晓敏


报告摘要:Gradient-based methods for the distributed training of residual networks (ResNets) typically require a forward pass of the input data, followed by back-propagating the error gradient to update model parameters, which becomes time-consuming as the network goes deeper. To break the algorithmic locking and exploit synchronous module parallelism in both the forward and backward modes, auxiliary-variable methods have attracted much interest lately but suffer from significant communication overhead and lack of data augmentation. In this work, a novel joint learning framework for training realistic ResNets across multiple compute devices is established by trading off the storage and recomputation of external auxiliary variables. More specifically, the input data of each independent processor is generated from its low-capacity auxiliary network (AuxNet), which permits the use of data augmentation and realizes forward unlocking. The backward passes are then executed in parallel, each with a local loss function that originates from the penalty or augmented Lagrangian (AL) methods. Finally, the proposed AuxNet is employed to reproduce the updated auxiliary variables through an end-to-end training process. We demonstrate the effectiveness of our methods on ResNets and WideResNets across CIFAR10, CIFAR-100, and ImageNet datasets, achieving speedup over the traditional layer-serial training method while maintaining comparable testing accuracy.


上一条:Seminar第2189讲 量子群与q-舒尔代数(ON QUANTUM GROUPS ANDq-SCHUR ALGEBRAS)

下一条:Seminar第2187讲 目标导向型减基法加速的多项式逼近法

  版权所有 © 上海大学   沪ICP备09014157   沪公网安备31009102000049号  地址:上海市宝山区上大路99号    邮编:200444   电话查询
 技术支持:上海大学信息化工作办公室   联系我们