you might be the first person talking about the data over the architecture! 🥲
we consider ourselves a data research lab! the vast vast vast majority of research was on making data that is truly general (ala a cognitive core) and 100% of our data is synthetic (but not the type of crap that is just spit out from an LLM obviously)
we built the agi compiler
it watches your llm agent work, finds the parts that are secretly deterministic, and compiles them into verified binaries that cost nothing to run
llms are just the first frontend, world models and new model types plug into the same toolchain, that's why it's a compiler
same 300 tasks, same verified answers, 6.4x less money
paper + code at the end
EvoEmbedding
An evolvable embedding model that maintains a latent memory queue to generate dynamic representations for long-context retrieval, outperforming static specialists 3× its size.