๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Transluce
@TransluceAI
Open and scalable technology for understanding AI systems.
๊ฐ€์ž… October 2024
21 ํŒ”๋กœ์ž‰ ์ค‘    11.8K ํŒฌ
Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. ๐Ÿงต(1/)
๋” ๋ณด๊ธฐ