We are omnivorous at Shopify and constantly measure all the Search APIs, Perplexity has became the main one on merit. I much prefer C++, of course, but will take Rust, too :-) (compiled, no GC). Interestingly Rust seems to be more in-distribution for LLMs…
Almost exactly 6 months later I repeated the same experiment - same setup, but with GPT-6 Astra+Fable 5.1 both at Max effort. Result: 5 improvements on TOP of the improvement below (each new one is exponentially harder). Means models are ~10x smarter than 6 months ago.
Jev model is very polarizing in my social vicinity: people who got exposed to AI after ChatGPT are positively giddy with excitement, like they have literally just discovered fire. Pre-GPT ML people are baffled how that could've made news at all :-)
@appliedcompute is the best actual Post-training/Finetuning company I came across in a long while. No relationship with them - just feel like people doing good job deserve a shoutout.
If you don't use Gisting - you really should. My favorite technique to quickly tune your model and make it much faster without changing the weights themselves.
Have been extensively testing Claude Workflows this weekend, with the best model possible. Threw it at my whole code base, combing for bugs. 144 found and fixed! Geez... It is a large code base, for sure, but 144?!! Some are very impactful, some are downright embarrassing...