Agent memory, from the team that shipped it:
"you don't want the latency of including everything"
here's the memory pipeline, step by step:
step 1 → 5:50 – working memory first – it has to hold this conversation before it holds a history
step 2 → 12:09 – extract memories from the conversation, not from the log
step 3 → 13:30 – "my name is Kim" said twice is not two memories
step 4 → 13:42 – consolidation – contradictory memories are worse than none
step 5 → 19:14 – generate memory in the background, never on the user's latency path
most people dump the whole transcript into context and call it memory
that's a log, and it gets slower every single run
watch & bookmark → then read the full graph engineering guide below ↓