注册并分享邀请链接,可获得视频播放与邀请奖励。

Horace He
@cHHillee
@thinkymachines Formerly @PyTorch "My learning style is Horace twitter threads" - @typedfemale
加入 February 2010
612 正在关注    53.4K 粉丝
Imo a lot of people don't think about the a2a cost in EP correctly. Andrew Gu explained this perspective to me a while ago and I think it's the right one.
If you calculate out comms the dispatch takes time input_dim*comm_time_per_byte and the matmul takes time input_dim*intermediate_dim*time_per_flop. If you divide these, you find that the ratio that determines whether you can overlap compute with comms is comm_time_per_byte/(intermediate_dim*time_per_flop). Since latent_moe shrinks input_dim by 2 but keeps intermediate_dim the same, there's no impact on whether you're network bound.
显示更多
0
5
149
4
转发到社区