New Tesla FSD technical talk.
For the first time, we learn fascinating details of how they train their end to end AI model and it turns out it’s exactly how humans learn.
When we first learn any complex task like driving, our conscious mind (cortex) is fully focused on all the little details our instructor has told us to focus on. Are you centered in your lane? Look out for stop signs. Pay attention to stop lights.
We are so mentally overloaded that we don’t drive very well. But what we are doing through cortex guided practice is training our unconscious brain areas (basal ganglia and cerebellum) to drive without our conscious cortex having to pay attention to and direct everything.
After many hours of practice we have effectively coded an end to end neural network in our brains that allow us to drive without much conscious thought at all.
This is exactly how Tesla trains their end to end FSD v14. In the data center, they use labeled training videos which tells the model being trained to pay attention to lane markings, stop signs and street lights.
But the resultant created model doesn’t have any explicit sections that look for these cues, it is all implicitly baked in the single end to end model that the car uses to drive.
Before FSD 12, the AI running in the car did have a 100 different feature detectors, which is why it drove like a brand new driver. It turns out that none of that engineering work was wasted, they still use it, but only when training a new model from scratch.
And yes, this has direct applicability to training Optimus.
Much more info in the talk, including how VLAs fit into this picture.
Start the video around 6m50s to skip stuff you already know.