Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied!
We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
Show more
Proud to partner with
@baseten and
@baselabs to build safety infrastructure for open-source models—which are essential to lots of safety research, including our own. Safety must be built into open models and provided by those who serve them, and we’re excited to help enable that!
Show more
LLMs are like Schrödinger’s cat: many possible trajectories, but you only see one outcome per run.
To really understand models, and debug where they go wrong, you can find the "forking tokens" that lead to different trajectories. Our new research does this 100x more efficiently!
Show more
We’re giving out $1M in grants of free Silico usage for academic and nonprofit researchers focused on AI interpretability and alignment.
We feel extreme urgency about advancing interpretability for alignment, and we want to help more researchers push it forward. 🧵
Show more
With just one prompt, we taught an LLM to see - then looked at its representations to debug where it fell short.
Qwen 3 8B can read text, but has no way to see or understand images. Silico trained a vision adapter that matches the official Qwen 3 VL 8B in multiple benchmarks. 🧵
Show more
Silico, the platform for ambitious AI research, is publicly available today.
AI is advancing fast. The tools to understand it need to advance even faster. Silico lets you interpret and train your models at frontier scale.
Learn more + get access 🧵
Show more
PCA reveals beautiful geometry inside a protein language model. But does the model actually use it?
I explored how protein folds—like this beta-propeller fold—are represented in ESMC-6B using the interpretability tools built into Silico,
@GoodfireAI's research platform 🧵
Show more
we're hiring a lead for model training & training research. if you want to make models great at interpretability, and push the frontiers of intentional design, apply here:
My thoughts after daily driving Silico:
It makes research substantially more joyful and exciting, allowing me to accomplish more and explore a more diverse set of methods and ideas. It's sticky; I don’t want to go back to not using it. The agents have more freedom and agency than in Claude Science. Just like how the abstraction from chat-to-agent is qualitative, agent-to-Silico feels qualitative because of the ability to rapidly explore many paths without needing to help the models much. It’s like speedrunning growing a bonsai tree, extending branches, pruning others. I tried Silico on both genomics and mechanistic interpretability. Also, tell me what I should try next
Show more
The folks at Pangram used theirs and produced a paper from it for a question I posed and it honestly answered the question better than anything I'd have come up with even after several days' work
Silico is a virtual lab for AI research. It designs experiments, manages GPUs, and presents results in a beautiful way. We believe it’s the future of AI research and want to enable everyone to study and understand the nature of intelligence.
Show more
silico is really great! if you do ai research you’ll probably find it a big uplift over cc/codex, especially for interpretability
Our team spent months developing RLFR, our method which uses probes on a model's internals as reward signals for RL.
Silico reproduced it in 2 days, reducing hallucinations in Qwen3-8B by 37% without capability loss. (3/6)
Show more
just welcomed chris earls, cornell engineering professor, to the team!
i'm particularly excited about his work to unlock scientific creativity in frontier AI with interpretability
we've hired several professors at
@GoodfireAI because we're investing heavily in foundational research to discover the science of neural networks. if this is work that you're interested in, join us!
Show more
Goodfire's Silico decided to show me this result of how representation of days of the week emerges over pretraining
If models think in shapes, our tools should too.
Our latest research: Block-Sparse Featurizers (BSFs), a new way to find concepts in model activations - using multidimensional “blocks” instead of single directions. (1/9)
Show more
Neural networks do math by rotating shapes.
We found a shape-rotating calculator hidden inside an LLM – and it’s used for more than just math! (1/6)
A simple example: days of the week, which lie on a circular path in models’ activations.
Steering linearly from Monday to Friday gets you incoherent outputs in between. Steering along the circular manifold means you cleanly shift from Mon → Tues → Wed → Thurs → Fri. (5/8)
Show more
Neural networks might speak English, but they think in shapes.
Understanding their rich *neural geometry* is key to understanding how they work – and to debugging and controlling them with precision.
Starting today, we’re releasing a series of posts on this research agenda. 🧵
Show more