注册并分享邀请链接,可获得视频播放与邀请奖励。

JJ
@JosephJacks_
加入 August 2013
540 正在关注    45.9K 粉丝
20 million ~ H100 GPUs equivalent of compute exists globally. 5% ~ of this training for 6 mo roughly produces a 10 trillion parameter SOTA model.. Astra / Fable level. The vast majority of compute is used for inference.. pre-training is billions of times less efficient than biology.
显示更多
0
11
591
12
转发到社区