๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Stewart Slocum
@stewpervised
prev AI alignment @xai, phd @MIT
๊ฐ€์ž… September 2019
245 ํŒ”๋กœ์ž‰ ์ค‘    1.4K ํŒฌ
Could existing alignment audits have caught the behaviors in the OpenAI-HuggingFace incident before they happened? As a first step, we reproduced the incident with public models. ๐Ÿงต
๋” ๋ณด๊ธฐ