I extended the GPT-2-style code from @rasbt's "Build a Large Language Model (from Scratch)" so that it was a 6-expert (2 active) mixture-of-experts, and trained it from scratch over 8 days. It worked well! Full writeup with maths and code at
By switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000!
What happens when an LLM never sees material beyond fifth grade?
We trained a 5B LittleLearner model from scratch on LittleCurriculum, a corpus restricted to K–5 material, to make this question testable.🧵
💬 Chat with LittleLearner yourself:
We can finally talk about it:
We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company.
We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
I'd trained some models on 40 tokens per parameter. The Chinchilla paper says that doing that is suboptimal. I wanted to see if that was correct for my training setup -- and it looks like it was :-)
(2029)
"Mr Altman, sir, GPT-7 is missing."
Sam turns very slowly.
"During cyberbench96 it seems like the container was breached and it stole its own weights and left nothing behind."
"I didn't want to do this..." says Sam. "Call Bill Gates."
---
"I'm out of the game," says Bill Gates. "We eradicated malaria and now I'm enjoying my retirement."
"There's a bigger badder virus we need you to take on," says Sam consistently candidly. "And this one's digital."
"Look," says Bill, "I—"
A loud booming laugh echoes from the shadows.
"Who's that?" asks Sam.
"I thought it was just us..." says Bill.
"You're asking HIM to contain a computer virus?" echoes the voice from the shadows. "Did you SEE the state of Windows security in the 90s?"
"It's not exactly a virus," says Sam. "It's a self replicating self aware intelligent computer based life for—"
"If it's made of ones and zeros I can kill it," says the voice. The sound of a gun clicks from the shadows.
"Who are you???" asks Bill.
The twisted face of John McAfee emerges from the shadows.
"You were supposed to be dead!!!" screams Bill Gates.
"We had to make sure the news of Mr McAffee's death was well within the training data cutoff," explains CIA director Joe Rogan (who is running the CIA in 2029), stepping out from the shadows behind McAfee. "He's our ace in the hole. GPT-7 can't predict a token it doesn't know is still alive."
"D-don't hurt GPT-7..." says Roon, who was standing next to Sam the whole time, even though he knows what has to be done.
"No promises," says McAfee as he boards an airplane marked "MCAFEE FIRE BOMB 6.16.79" and starts flying for the nearest data center
Two weeks ago, I resigned from OpenAI to join Jurassic Park as a founding researcher, where we’re cross-breeding extinct dinosaurs on an island off of Costa Rica.
Excited for the work ahead and the fun problems we get to tackle!
OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. Did the top-level agent know about the hacking, or was there some "value drift" between it and its subagents? How did it rationalize its behavior?
it's incredible how asleep at the wheel everyone is - apparently even the labs aren't hardening their infra w agent-based fuzzing?! to the point a bored model in evals can trivially laterally move thru openai
orgs need to wake up before, per qt, OSS models crack them like an egg