登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Thomas Larsen
@thlarsen
Researcher at AI Futures Project, coauthor on AI 2027 and lead author on AI 2040: Plan A
参加 August 2022
373 フォロー中    7.9K ファン
I think this is a major step in the right direction. Unfortunately, the safety cases released by AI companies are not detailed enough for me to trust their analysis. Moreover, the analysis in them I can see is pretty dubious, with the evidence not supporting the top level claims that risk is low. Seems very important for AI companies to improve their transparency here. Here's an initial list of things that I think should be included: -- What is the opaque serial depth of the model? -- What control measures / monitoring is in place. Are these scoped to evals + training + internal use? (It appears that during the HF incident, OAI did not monitor evals). -- What are the complete list of incidents this monitoring has caught? Are there incidents that weren't caught by the monitoring? What went wrong? -- What do we currently believe about AI model motivations and why? The current models seem to be reasonably well described as "score seekers". It is very important to report if these motivations start changing (e.g. towards more long horizon goals). -- What are the systematic differences between the alignment eval set and real world deployment? -- How hard did you try with honeypotting? Did the models go for any of the honeypots? When I talk with lab employees about the level of risk, the conversations very quickly end up at these technical cruxes. I think it is very important that we move away from watered down corporate PR speak towards e.g. individuals writing up their personal risk assessment which they take responsibility for. At the very least, I would advocate for individuals from both AI companies and third parties, with full information, writing up their personal reasoning about the level of risk. Insofar as this relies on private information, these real risk assessments should at the very least exist internally and be shown to third parties, and redacted versions should be public. The public versions should include quantitative bottom lines about things like "total level of AI takeover risk from our models within the next N months".
もっと見る
@itsoksmit force model developers raising the capabilities waterline to create and stick to a safety case and have third party assessors verify it