註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Samuel Shvartsman
@SamuelSBlackman
RL/Robotics @wendylabsinc
加入 May 2026
472 正在關注    181 粉絲
Day 12 of making an autonomous G1 bartender. First thing is first, the policy on the real G1 is now 30s in total, just like in sim. However, our sim policy is still not optimal and will not pick up all cans. Because we dont have much time and need to get this working in the next 5 days and RL will just not work if hyperparams are off, We decided to rent out a bunch of ~25 5090s (they are under a dollar each) and run a massive number of Protein (from pufferlib) sweep variations. We manually tested out simple things like changing the reward values from completing an entire task and we got completely different success rates within 24 hours. The default rewards that astra gave us were not ideal. So, with this Protein sweep, we have a single 5090 as an orchestrator and 4 other 5090s computing different starting values and this gives us 5 batches of sweeps. I have decided that treating the reward numbers could also improve the training, so by tmrw, I really hope that we get good hyperparameters so that my sanity could remain :)
顯示更多