๐ Context Engineering in 2026: Compaction, Memory & Cost
@Whats_AI,
@samridhivaid and
@omar_solano1 return!
This workshop is about engineering the context window so rot stops happening, shown with
@towards_AI's open-source AI tutor, which answers questions for students of our AI-engineering courses.
Context engineering is deciding what the model sees on every single call โ instructions, history, retrieved course content, memory, and tool outputs โ and it's the line between a tutor that holds a coherent session and one that forgets the student's setup halfway through.
We'll move in three stages, mirroring how the project actually went. The concepts:
- the two root problems (a finite window, a stateless model),
- the full compaction toolkit (truncation, trimming, tool-result clearing, summarization, and offloading to files โ and when each actually helps),
- memory that survives across sessions, skills loaded on demand, and
- production-grade retrieval (chunking, metadata, course scoping, hybrid search, reranking, and evaluating).
We'll cover the tutor's architecture, and the evaluation harness we used to measure every run on Gemini โ tokens, cost, latency, and memory probes instead of vibe-checks. At real volume, even Gemini Flash got expensive, so we tested whether open and local models could match the quality for a fraction of the cost and match result quality.
Everything is open-source and will be shared during the workshop.