backdrop: kings all asleep/on vacation while minions slave away
august 2024 - first big fail
jan 2025 - only competent vp bails after one terrorist lead demanding an even bigger slice of pie breaks the camels back
summer 2025 - second big fail + the king of terrorists climbs the ranks above the slumbering execs
summer 2026 - slumbering execs + terrorists who only show up for successful announcements move on, leaving just the king of terrorists behind
-----
lesson: if you're high enough up there, you can pretty much be a terrorist and hold orgs hostage while continuing to ask for more and more and more, until there's absolutely no more aura or gold left to plunder, and then just move on to greener pastures to do it all over again!
(honestly, good timing, the king of terrorists is just insufferable, no?)
Show more
the reason "mid-training" fails to sustain its own paradigm of research at some of these major labs comes down to individual incentives more than scientific ones (despite how much people claim it to be irrelevant in some fully e2e codesign-with-applications utopia).
to explain the dynamics here, we have to first look at how "post-training" developed as a separate entity: pre-training and scaling these models historically have been more of an infra problem coupled with "rigorous scaling laws" to justify compute-time investment that exponentially increases YoY.
this type of scale and complexity attracted more of the "purists", aka people who were relatively early to buy into the scaling paradigm and early to language modeling in general.
(the modeling domains of vision, robotics, speech, etc historically have not required similar levels of multi-data-center level scale to be useful for practical applications, though they certainly have / continue to have their own individual scaling-up-to-utility curves)
by the time LLMs took off, interest to work in this space also grew exponentially, but these "core problems" of scaling LLMs didn't similarly grow in complexity to absorb this new interest.
furthermore, there was still a very real, very significant gap in bridging the raw LLMs to domains/applications, so "post-training" developed both as a way of organizing people alongside the final deployment surface area, in order for an R&D artifact to successfully make its way into actual use-cases that made $$ or hype.
(as an aside, "post-training" is also where you see the most churn, since it involves 10x-100x more people, each either aiming for bits of the jagged frontier or getting locked out of bigger pools of compute for bigger scaling efforts)
the unfortunate bit here is that as the post-train surface and sprawl increase, so does the disconnect with pretraining decisions that are made in a vacuum (relatively speaking). in theory, this can be bridged if coherent metrics and signals are backproppable into pretraining decisions, but that endeavor similarly suffers from practical resource/execution constraints.
so again, in theory, there is ample room for a "mid-training" space to grow to fill this judgement/signal gap. however, if your organization incentivizes individual-hero-promo-maxxers, what this turns out to be is a turf-war of getting squeezed out by both pretraining and post-training teams. in order for you to succeed, you need to lube up two major political organizations, both of whom view you to be a threat to their zero-sum land of finite impacc/compute allocations (and more concerningly, a threat against a saturation of their own success metrics that have all been hill-climbed into near irrelevance).
so at some point, if you can deliver value in this space, you'll take a hard look at these challenges and wonder whether the upside of success is worth fighting these largely non-technical political battles wrapped over actual technical challenges. either way, with enough time and reorgs, you'll simply get absorbed into pre-/post-train orgs anyway, as a duality of power is much easier to balance at an exec level, compared to settling squabbles amongst a triumvirate.
Show more
chinese is peak token efficiency, babes
the nytimes really didn't hold back on getting internal chat messages from anthropic
where the "same people" who previously claimed the ability to bring about a cybersecurity "reckoning"
are now left wondering if they're being "picked on, bullied, unfairly targeted" by the administration
and are also leaking these chats to the press
-----
nothing like some tough times to see who abandons ship first
Show more
there is no better time in tech than now to be a jack of all trades, master of a few.
just make sure to keep adding to the few year over year, such that the cumulative breadth of expertise you collect becomes an increasingly rare combo. remember, if you're top 10% in 3 different areas, that already makes you top 0.1%. keep switching it up until you get to "your best", and then switch it up again (great for a particular flavor of people who don't enjoy resting on laurels, maybe not so great for others).
question all institutional value and pedigrees, all traditional career paths or corporate ladders: the college industrial complex is getting shaken up, alongside a disappearing managerial class, so if you're pursuing either make sure you are fully internally aligned with why. social/political capital in a particular institution can feel incredible, but if you're spending all your energy on complex political people games, you're not a technologist anymore, you're an unelected politician. if you're ok with that, then all's well.
critical thinking is more important than ever: take nothing at face-value, question everything and everyone. the equivalent of ai slop can be found in humans operating under misaligned incentives and interests. the sooner you're clued into disambiguating the talkers/larpers from the doers, the better off you'll be figuring out where and who to invest your time in.
the anxiety of job displacement is very real, since a surprising amount of white collar work/prestige is built on a performative house of cards, significantly lacking in correlation with technical breadth, depth, and skill. as long as you keep learning, keep building, keep producing receipts, you will be fine.
if all that sounds ok to you, welcome to the world of technology! it's truly one of the few places you can experience child-like wonder every few years, and be constantly humbled & excited by new adventures, as scary as they may seem at first.
don't give up, drink your water, get your sunlight, and take breaks as needed. tech careers are notoriously nonlinear, so you might as well embrace it and enjoy the ride!
Show more
in the (human-observable) limit, these language models will all be speaking chinglish:
- most highly-educated digital workers speak english or chinese
- peak token efficiency <> expressiveness
- compact translations of those four-word chinese idioms (成语) or continued language compaction with gen-z slang and beyond
- common grammatical structure amplified between the two languages when expressing concepts in corporate environments (this++ if/when there's also overlap with hindi)
or maybe more simply: we're purely in a digital data numbers game (against a perfect ai translation amplifier) from here on out
Show more
Incredible writeup! Some notable 💎s:
Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens.
They also use separate attention modes for disaggregated prefill vs decode (is this the first public account of arch difference between the two? 👀).
1/🧵
Show more