superagent_

blog

thoughts, updates, and insights from the superagent team.

research·July 6, 2026·4 min read

When Terminal Output Owns Your Clipboard: OSC 52 in Warp

Affected Warp builds honored OSC 52 clipboard escape sequences from terminal output, allowing silent clipboard reads and writes with no default-deny gate.

read more

research·March 24, 2026·5 min read

Frontier models miss 57% of threats in agent context

We ran 485 real artifacts through Claude 4.6 Opus with a security-focused system prompt. The model missed 57% of the threats brin had already identified. Here's the full breakdown.

read more

research·January 21, 2026·3 min read

We Bypassed Grok Imagine's NSFW Filters With Artistic Framing

Text-to-image safety is broken. We generated explicit content of a real person using basic compositional tricks. Here's what we found, why it worked, and what this means for AI safety systems.

read more

research·January 13, 2026·5 min read

The Threat Model for Coding Agents is Backwards

Most people think about AI security wrong. They imagine a user trying to jailbreak the model. With coding agents, the user is the victim, not the attacker.

read more

research·November 19, 2025·2 min read

AI Is Getting Better at Everything—Including Being Exploited

As AI models become more capable and obedient, safety improvements struggle to keep pace. The GPT-5.1 safety score drop reveals a structural problem: capability and attack surface scale faster than safety.

read more

research·November 17, 2025·5 min read

Are AI Models Getting Safer? A Data-Driven Look at GPT vs Claude Over Time

Are frontier models actually getting safer to deploy—or just smarter at getting around guardrails? We analyze 18 months of Lamb-Bench safety scores for GPT and Claude models.

read more

[ ← prev ]12[ next → ]

join our newsletter

updates on securing code and agents, vulnerability research, and product news.