.@EvaLongoria will present Thalía with the Icon Award at Billboard Women in Music 2026 🏆👑
Tune in LIVE on Wednesday, April 29 at and starting at 6:30 p.m. PT.
Institutions evaluate real corporate risk before committing to shared infrastructure, including who controls it.
Canton was built with that in mind.
@YuvalRooz at @Decasonic's Web3 Investor Day on Canton's governance structure.
@hud_evals@ycombinator Tera — zero-token browser use
solve once, reuse forever: a browser RL env that learns from past traces, then replays matched workflows with 0 LLM tokens.
took #3# overall.
@hud_evals@ycombinator Warehouse AI
an RL environment for autonomous warehouse robot fleets: fulfill orders, avoid collisions, coordinate dispatch, and learn before touching the real world.
As eval is downstream of everything, it determines whether you will spend your time optimizing the right metrics.
The current gap between academia and industry AI labs is the attitude toward eval.
In academia, the eval set is very hard to change since a) you need to explain why your eval is better and b) you need to benchmark against your cited works with the new eval and show that your work is superior.
Doing both a and b at the same time invites risky rebuttal, even if you are doing a good job on a. It is far easier to benchmark against the eval set that everyone has agreed upon.
In contrast, in industry AI labs, customer feedback is your eval set and it keeps changing to cover the long tail that you could never think of during years of PhD programs.
If the loss functions are not a good proxy for customer feedback, then you change them until both are aligned.
Thus, academia might train students who are very good at hill climbing but inexperienced in building eval sets that capture hard real cases. To move the needle, building the right eval set matters the most.
In evals, Sonnet with an Opus advisor scored 2.7 percentage points higher on SWE-bench Multilingual than Sonnet alone, while costing 11.9% less per task.