We’re excited to announce that Trishool’s HaloGuard 1.0 𝐡𝐚𝐬 𝐚𝐜𝐡𝐢𝐞𝐯𝐞𝐝 𝐒𝐎𝐓𝐀 prompt-safety performance among open-weight guard models.
Today, we present HaloGuard 1.0, a constitutional input classifier for multilingual AI safety.
It is built as a first-layer input guard that checks user prompts before they reach a downstream LLM, agent, or application.
This is part of the safety infrastructure being built through
@trishoolai , our decentralised AI red-teaming subnet on Bittensor SN23.
Full arXiv paper goes live soon.