Cool new post with Lily tracking where LLM values come from.
My takeaways: SFT is (again) a big deal for model behavior, and there’s lots of important and low hanging post-training science still to do!
New blog post! LLMs exhibit value preferences (e.g. intellectual integrity🤓, warmth🥰) which affect their responses to subjective user queries. These change during post-training, sometimes unexpectedly. We ask: can we predict these value changes from just training data? Maybe!🧵