continuous learning at 2 billion parameters on a single device with a custom personal-planning env
this is a fun way to tinker with continuous learning problems. I've setup continuous calendar env.
a small model plans a synthetic day, gets interrupted by a meeting, and has to repair the plan without reshuffling everything.
the current setup:
- OpenEnv for the environment. Costomised and preference based version of CalendarGym
- explicit preferences and rule-based scoring instead of an LLM judge
- SFT on corrected choices, mixing older examples into later training rounds
- evaluate on the same held-out days after each round
after 3 rounds: half as many simulated correction flags than a frozen model with access to the same memory pool, across 384 test days and 3 runs.
the best result is SFT + replay, still can't get SDPO to work. a flag means another offered option scored better, not that a human had to intervene.
still a toy, but now it’s a personal-planning bench i can just iterate on.
you can run the environment locally and plug in your own policy: see below