Register and share your invite link to earn from video plays and referrals.

Lawrence Jang
@JangLawrenceK
@mldcmu phd student, @siagents
369 Following    514 Followers
Luna is one of my favorite models out right now, it also does really well on the computer-use benchmarks we tested On Odysseys ( and MyPCBench ( it scores 51% and 55.4%. To put this into perspective, it performs similarly to the heaviest models offered while being 20x+ cheaper. If you have the max codex plan, this is essentially unlimited usage. I am super happy about this level of model being offered at an even cheaper rate and would suggest try asking Luna to use computer-use in your workflows
Show more
guys I just cancelled my Claude plan I don’t know what happened
An annoyingly common question I get as an AI PhD student is “When can you get the ChatGPT AI to do something useful? It can’t even work on my phone yet. Siri is pretty dumb.” To be fair, I think their criticism is correct. I personally wished that AI was better integrated on my phone. LLMs can solve IMO problems, so shouldn’t it be a cakewalk for it to remind me of the text I forgot to respond to last week? Obviously not, since it doesn’t exist in my pocket yet. Or maybe Apple’s new update yesterday fixed this and my research project is obsolete. We are releasing iOSWorld ( a dynamic iPhone benchmark with 26 newly created apps grounded in personal context. Each of the 26 apps is centrally seeded around one persona, Jordan Avery, and the apps interact together in a realistic ecosystem that reflect real app interactions. We create 133 personalized mobile agent tasks to test in this environment, and the best model, even with privileged information, only scores 51%.
Show more