Your AI agent's memory is probably stuck inside one app.
Walrus Memory (@WalrusProtocol) fixes that with an MCP that creates memories as you chat, then can recall them anywhere - Claude Code, Pi, etc.
Portable and actually owned by you!
Karpathy's LLM compiler idea: raw docs -> LLM -> wiki articles -> agent queries at runtime.
Everyone's applying it to external data.
But your Claude Code session logs are the same thing. Raw folder already exists. Nobody's compiled it yet...๐ค
I just started a Discord for anyone who likes gaming, hanging out, and non-toxic lobbies ๐ฎ
Still small right now but weโre building a fun community.
Come join us ๐พ
SWE-bench Verified is contaminated. OpenAI just published the proof.
Top models: 70%+ on Verified, ~23% on SWE-bench Pro.
All frontier models can reproduce original fixes from memory. 59% of hard tasks have flawed tests.
The benchmark everyone was citing? Meaningless now.
Best coding agent tip I've tried recently: just add "use red/green TDD" to your prompts. Agent writes tests first, confirms they fail, then implements. Catches broken code and unnecessary code. Two words, pretty big difference in output quality.
Microsoft confirmed a Copilot bug has been summarizing confidential emails since late January, bypassing DLP policies.
Not a prompt injection. Not an attack. Just a bug that gave AI access to data it shouldn't have processed.
Every AI assistant that goes deeper into your data increases the blast radius of a single misconfiguration.