Grok 4.6 was the first model trained on internal model-development tasks.
My team built the training and evaluation stack that enabled Grok to learn from work that accelerates model development itself, including production inference and kernel optimization.
Grok 4.6 now leads our internal MTS Eval and InferenceEval. In one experiment, we gave an earlier checkpoint our production inference codebase in a fully automated environment. It explored 297 optimization ideas and shipped 3 changes on top of an already human-optimized stack, improving prefill throughput by 3.1% and decode throughput by 1.5%.
Our work is featured in the R&D Enablement section of the Grok 4.6 model card.