“Insilico Medicine will contribute proprietary datasets spanning multiple specialized life science tasks and execute closed-loop wet-lab experimental validations on the outputs generated by the foundation models
during training.”
pretty big deal - Data is the big blocker to making a model superhuman in a domain
as AI solves the world’s hardest problems, data curation requires doing things in the real world (ie. Wet labs)
will probably see a lot more of this style of deal across domains
in one way it’s incredibly exciting to have collaboration to solve humanities hardest problems in Bio, Physics, Math, Robotics, etc
but also hope some of this data becomes open
it’s tough game theory for openness
in a domain, say there’s N co’s with valuable data. Labs can broker a deal with any few of them for data, access to their models (see Rosalind though not exactly the same), and revenue share
Labs can also purchase any of those companies themselves
And Labs have massive data + compute moats
but clearly from the last few months of open data + training research, the open community is getting very strong at building near the frontier and only scaling