been locked in trying to set up a crude continual learning environment that I can develop into a personal continuous learning bench.
nothing works yet, but here's the plan:
- take some daily tasks that agents can do but lend themself to customisation, like daily planning.
- build an rl environment around the daily tasks
- add synthetic preferences to the environment. i.e. how i like my day planned
- get a judge model to judge traces based on preferences
- get agents to do the tasks and record the traces as buckets
- benchmark the base model on the task (LiquidAI/LFM2.5-1.2B-Instruct)
- SFT the model on the traces
- SDPO a model on the traces with a judge model hint
So far, Evals work, SFT works, but SDPO doesn't. Pretty sure that the preferences are too arbitrary, so considering manually adding more preferences to setup a better few shot judge.