Register and share your invite link to earn from video plays and referrals.

Search results for PromptInjection
PromptInjection community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including PromptInjection
GPT-6 Astra has amazing prompt injection robustness. Just in time for for how we're all using (or want to use) AI
turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn't expect that a year ago. auto mode is default in claude code as of next week
Show more
0
188
2.8K
163
Forward to community
Favorite new AI phrase: benevolent prompt injection Definition: The act of assembling a packet of context automatically, across the agent environments our engineers use. Managed startup hooks inject a project's rules, requirements, process, and current state before a new task begins, so a new session does not depend on a user starting with the right prompt. Killer concept by @tenex_labs' engineers.
Show more
Surprised we haven’t seen more publicly disclosed email prompt injection attacks, especially with consumers running a Claw or equivalent. Am I too paranoid to think email access should only be given to a model with no other tool / internet use, with labeling and drafts via Gmail MCP as the only allowed actions?
Show more
An AI agent that holds an API key is one prompt injection away from using it. Hardening the sandbox raises the cost of escape — but it doesn't change what's inside. New on the TRM Tech Blog, Preet Bhinder digs into the credential broker pattern that keeps secrets out of agents' hands entirely, and the four-part process for shipping security-critical code with AI in the loop. Read the post here 👉
Show more
# Practical ways to use the Claude Agent SDK 🛡 Build defense in depth with prompt injection countermeasures, credential proxies, and least privilege. Secure Deployment provides isolation techniques, credential proxy patterns, and network/filesystem controls for safe production agent operation. 📌 Title: Deploying AI Agents Securely 🔗 URL: 🧩 Overview Combines sandbox-runtime / Docker / gVisor / Firecracker isolation, credential proxy patterns, and network/filesystem controls for multi-layered defense. 🛠 How to use it Choose isolation by threat level: sandbox-runtime for single developer/CI, gVisor/Firecracker for multi-tenant or untrusted content. Inject credentials via Envoy / mitmproxy / LiteLLM proxies outside the agent boundary. 🏗 Practical usage - Credential proxy pattern: inject API keys at a proxy outside the agent boundary. The agent calls APIs without ever seeing credentials. Route via `ANTHROPIC_BASE_URL`. - Least privilege: mount only required directories as read-only, exclude `.env` / `~/.aws/credentials` / `*.pem`, use `--network none` + Unix socket proxy. - In cloud environments, use private subnets + cloud firewalls to block all outbound except through the proxy, which enforces allow-lists, injects credentials, and logs all traffic. 💡 Use cases 🔐 Credential management outside agent boundaries 🌐 Network restrictions preventing unauthorized data exfiltration 🏢 Strong isolation for multi-tenant environments ⚠️ Watch out Untrusted content (READMEs, web pages, user input) may contain prompt injection attempts. Network controls are your last line of defense — always configure them. #ClaudeAgentSDK# #AI#
Show more
$META is also putting Muse under a public bug bounty program Researchers can earn up to $300K for critical security or prompt-injection flaws, including $250K for a fleetwide Muse compromise and $130K for compromising a single user’s agent.
Show more
😱 Is your LLM truly bulletproof? 🛡️ A single prompt injection can make your AI break its own rules. Attackers are constantly testing your defenses — is your system ready? Don’t wait for data leaks or AI failures to happen! ⏰ July 9 | 20:00 (UTC+8) Join Mason, former AI Algorithm Expert at ByteDance 👨‍💻 Learn how to break, test, and secure LLMs through real-world attack & defense scenarios! 🎁 Get the TinTinAI resource pack + lucky draw rewards during the livestream! 📌 Wechat Video Channel [TinTinAI] ⬇️ Scan & Reserve your spot #TinTinLand# #TinTinAI#
Show more
Using Fable and first time I’ve seen a proactive notification of an agent flagging and ignoring a prompt injection attack. Pretty cool!