註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

JJ
@JosephJacks_
加入 August 2013
540 正在關注    45.9K 粉絲
20 million ~ H100 GPUs equivalent of compute exists globally. 5% ~ of this training for 6 mo roughly produces a 10 trillion parameter SOTA model.. Astra / Fable level. The vast majority of compute is used for inference.. pre-training is billions of times less efficient than biology.
顯示更多
0
11
591
12
轉發到社區