13/ What I take away from all this is that all the training to create the model is engineered around crafting the best possible ICL mechanism. Further training degrades this mechanism, at least for knowledge acquisition, and maybe continual learning should just focus on loading the right information up for icl (ie into the context window)
So this provides an answer to the architecture question, at least for now: the channel worth engineering for memory is the one that comes with addresses ie the context.
Paper:
Code and data to follow shortly
Done at
@baseten.