Ken Thompson’s “Reflections on Trusting Trust” feels super relevant to AI.
He bootstrapped a “poisoned” compiler that left no traces of the poison in the source code because it compiled itself.
You can imagine the same thing happening with models where one poisoned generation helps train the next while removing any traces of the poison.
显示更多