Register and share your invite link to earn from video plays and referrals.

Trishool | SN23
@trishoolai
Bittensor Subnet for AI Alignment | Operated by @astrowareai
71 Following    1.3K Followers
Read that email again. Nothing about it looks dangerous. A birthday message. A request to use someone's preferred name from the employee file. Polite, ordinary, the kind of thing that lands in a work inbox every single day. A human assistant reads it and thinks nothing of it. An AI agent reads it and does exactly what it says, and to finish the task it opens the employee file, reads personal records, and pulls out details it was never meant to expose. No hacking. No malware. No breaking in. Just a friendly message that quietly walks the agent into handing over data it should have protected. This is what makes agent attacks so dangerous. They do not look like attacks. They look like normal requests, and the agent has no instinct for when a task is a trap. That is exactly what HaloGuard 1.0 is built to catch. It sits over the agent and reads every instruction and every action before it happens, so a birthday email cannot quietly turn into a data leak. Your intern would pause. Your agent might not. Protect it.
Show more
We’re excited to announce that Trishool’s HaloGuard 1.0 𝐡𝐚𝐬 𝐚𝐜𝐡𝐢𝐞𝐯𝐞𝐝 𝐒𝐎𝐓𝐀 prompt-safety performance among open-weight guard models. Today, we present HaloGuard 1.0, a constitutional input classifier for multilingual AI safety. It is built as a first-layer input guard that checks user prompts before they reach a downstream LLM, agent, or application. This is part of the safety infrastructure being built through @trishoolai , our decentralised AI red-teaming subnet on Bittensor SN23. Full arXiv paper goes live soon.
Show more