Looks very exciting! It always seemed likely that the optimal model architecture would be one forward pass per logical step. LLMs are quite wasteful on one end of the spectrum at one forward pass per token. This looks more like one forward pass per prompt?
Would be very interesting if one could somehow chain this together now into multi-step (i.e. multi forward pass) reasoning, and optionally in the end unpack cheaply with small LLM.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution