๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Eiso Kant
@eisokant
Co-CEO @poolsideai // President @ PIC // Board @axon_enterprise โ€œThe best way to predict the future is to invent it.โ€ - Alan Kay
๊ฐ€์ž… September 2007
4.3K ํŒ”๋กœ์ž‰ ์ค‘    13.4K ํŒฌ
Leaving image generation to diffusion but reasoning to an LLM is ๐Ÿคฉ Now take a step further and think about it in the context of robotics with spatial-action models. Work combining modalities is still in the early days.
๋” ๋ณด๊ธฐ
8/ With that, we reframed multimodal generation as structured text/code generation. Diffusion just renders pixels. Planning, logic, reasoning all live in the LLM โ€” so training looks like normal LLM training, and inherits all benefits of it: data + model scaling, reasoning, RL, tool use.
๋” ๋ณด๊ธฐ