Most graph-based memory stacks rely on an LLM to do the extraction work, but that gets expensive fast.
The extraction layer feeding an agent's memory should be cheap enough to run constantly as new data becomes available.
We're teaming up with
@cognee_ on September 3rd in Berlin to host an event on how smaller, local models can fit into that agent memory pipeline.
@m_newhaus, from our technical team, will demo how GLiNER handles entity and relation extraction for knowledge graphs, and
@tricalt will be talking about why running small models locally beats frontier LLMs on memory tasks.
🔗 to sign up: