Register and share your invite link to earn from video plays and referrals.

Kun Chen
@kunchenguid
Member of the Technical Community. Captain of @myfirstmate @theSSHHIP Author of
Joined September 2021
191 Following    37.4K Followers
my most important lessons on using voice input on agents talking with voice is auto regressive, like an LLM: - we speak one word at a time - we think and talk simultaneously - and we can’t easily take back what we said this created some interesting characteristics: - what we say first will influence what we say next. try starting your sentence with “the most important two things i want you to do are…” and you’ll find yourself cornered into a certain way of framing your message - we generally don’t have a ton of time to think while talking. and the faster we talk, the less thinking we do. this is why people who trained themselves to take pause are often more eloquent and perceived as thoughtful - we often perceive talking as higher stake than typing because correction is a lot harder these traits made voice input a terrible method for writing content. say you want to write a blog post, and using voice means you have to think through most things upfront, and randomly derail the content during talking due to choice of words, and can’t correct what’s written (unless you fallback to typing in the end) BUT.. when using agents, most of these traits don’t apply! agents can understand messy rambling just fine. they can understand your corrections if you said “no no scratch that, let me rephrase”. and they couldn’t care less if you take a pause so when using voice input on agents, you have to think of it differently from talking to real people or writing real content my advice: - don’t try to speak only the “final result”. speak your thinking process out loud this gives yourself a lot more room to reason through the concepts and arrive at something that actually represents what you want it also gives the agent more context on why you made certain choices you can even intentionally use phrases like “let’s think through this..”, or “here’s why..” to elicit more thinking - correct yourself often when you realized something you said earlier was wrong, you can either just stop the voice input and discard what you said so far, or start correcting yourself the good thing about agents is that they don’t judge. they are trained to never shame you for making mistakes or walking back your decisions once you started doing this more, you will perceive voice input as low stake, which makes it feel more natural and less demanding on the other hand, typing is text diffusion - you think in your head and type out the results, and you constantly refine what you already wrote it’s great for writing polished content, or precise text such as code. but it’s not as good for prompting agents because all your thinking tokens were hidden so the agent only sees the final “what” but rarely the “why”
Show more