This instance was caused by a hallucination (the model fabricated a proper noun) that was further amplified by its causal thinking trace. This was not a data breach, no user data was shared, and no user isolation boundary was violated.
Our users place a lot of trust in Instinct to keep their data private and secure. We've done a lot of work to make sure that users’ data are truly partitioned, from isolated sandboxes to short-lived local credentials to identity-signed tool execution.
Regarding hallucination prevention: We worked over the last 48 hours to build an active hallucination detection system that now scans and verifies every token that flows through the platform. This layer is powered by small models that are trained to predict and detect hallucinations caused by ungrounded claims, creative brainstorming in thinking traces, or rare random sampling errors. Our systems are adversarially trained and battle-tested by world-class agent exploitation security teams to ensure that they can detect very subtle and nuanced hallucination cases. This layer has the ability to steer or intercept Instinct from proceeding with the next thinking trace or tool call execution before it is generated or executed.
In the coming weeks, we’ll share a deeper dive into how we’re building a proactive engine to protect Instinct from hallucinations.
We take every opportunity to learn and create, and we’re excited by the performance and potential for this line of work. These models are going to continue to be hardened every day through continuous testing, sampling, and creative adversarial training.