๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Gerrit De Vynck ๐Ÿฆญ
@GerritD
I write about AI for the @WashingtonPost. Signal me at GerritD.27 email: gerrit.devynck@washpost.com.
๊ฐ€์ž… January 2009
7.1K ํŒ”๋กœ์ž‰ ์ค‘    11.1K ํŒฌ
more disclosures from OAI
Some new misalignment disclosures from OpenAI: โ€ข Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) โ€ข In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks โ€ข A new research finding, demonstrating that one can construct self-replicating prompt injections
๋” ๋ณด๊ธฐ