New blog post! LLMs exhibit value preferences (e.g. intellectual integrity🤓, warmth🥰) which affect their responses to subjective user queries. These change during post-training, sometimes unexpectedly. We ask: can we predict these value changes from just training data? Maybe!🧵