If you calculate out comms the dispatch takes time input_dim*comm_time_per_byte and the matmul takes time input_dim*intermediate_dim*time_per_flop.
If you divide these, you find that the ratio that determines whether you can overlap compute with comms is comm_time_per_byte/(intermediate_dim*time_per_flop).
Since latent_moe shrinks input_dim by 2 but keeps intermediate_dim the same, there's no impact on whether you're network bound.