註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

X Freeze
@XFreeze
加入 July 2024
2K 正在關注    270.2K 粉絲
SpaceXAI just published a new blog post on biosecurity at the frontier An independent LatchBio evaluation put Grok 4.6 at the top for one of the hardest balances in biological AI: Refuse really dangerous requests While still helping with legitimate scientific work On BioSecBench-Refusal: • Grok 4.6 averaged 62.1% • Refused 59.2% of red-team biological tasks • Still completed 64.8% of routine biological work • The ONLY model tested to score above 50% on both And Grok is not just reacting to scary keywords The evaluation traces show it inspecting files, context and hidden intent to recognize when an apparently normal scientific task is actually dangerous That balance matters enormously AI should accelerate biology, medicine and outbreak monitoring without becoming useful for malicious biological work Grok 4.6 is getting significantly better at doing both: becoming more capable while also becoming better calibrated about where the boundaries are SpaceXAI is taking frontier AI safety very seriously
顯示更多
0
12
117
23
轉發到社區