註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

nathan chen
@nathancgy4
learning, entropy-maximizing, opinions
加入 April 2022
730 正在關注    20.7K 粉絲
glm 5.3 flash uses the same kda gate lower bound, even the same numerical value, as kimi k3. k3's tech report was released exactly a month ago. if there's no coincidence here, getting such a model done within one month is quite insane (plus ox alpha was out a week ago..)
顯示更多