Register and share your invite link to earn from video plays and referrals.

ali
@aliteracy
your favorite robots' favorite robot @RhodaAI @Stanford
746 Following    1.4K Followers
this is a crazy response to very reasonable feedback on the paper LOL why must we optimize papers for clicks and engagement? isn’t the point of writing papers to disseminate knowledge into the community? this is seriously a bad look for all coauthors… fwiw, why not use SoTA linear layers like KDA or GDN2 for ablations?
Show more
"we found that distilling a model into an architecture that is very close performs much better than one that is very different to the teacher model."