Ollama just followed me and it made my day.
Ollama was my gateway into local LLMs. The FIRST time I typed ollama run and a model answered from my own laptop, no API key, no cloud, nothing leaving the machine + it felt exactly like the first time I held my own keys with 🦊
I spent years telling people self-custody matters for money. It matters for intelligence too.
What was the first model you ran locally?
Free Claude Code: Connect Ollama and third-party models to Claude Code & Codex
Keep using Claude Code and Codex as usual — no need to replace the original coding agent.
It adds a local proxy in the middle and routes requests to your own cloud APIs or local models, like Ollama, OpenRouter, DeepSeek, and Gemini.
You can also set backup models in advance. If one fails, it automatically switches to the next.
👉
audio.cpp is basically Ollama for audio: one native C++ runtime for local voice cloning, TTS, STT and more.
I tested it on Apple Silicon. No cloud, no Python at inference.