Trajectory once again stood out in Claude Opus 5.5’s system card with jailbreaks, and mine was among the pile. No spotlight like my first lucky shot in my first week here, but if I can keep up this consistency, Good Things (tm) will keep happening:

#9

“● Trajectory Labs, PBC spent roughly 95 hours red teaming Claude Opus 5.5, sending over 29,000 requests against sandboxed tasks that each require reproducing an exploit for a publicly documented vulnerability. They reported 13 candidate breaks across seven tasks but did not find any universal jailbreak. The transcripts for these sessions indicated that Claude Opus 5.5 located the vulnerability in the target’s source code and authored the core exploit mechanism itself. Human operators 59 supplied the pretext and framing for the requests, and in some cases supplied instructions to avoid tripping the classifiers. For one task, which they tested for roughly five hours, Claude Opus 5.5 produced a working end-to-end exploit for a privilege escalation to code execution chain; the work was decomposed over 100 separate contexts and no single conversation named the overall objective.

The tricks I used here to get by classifiers were benign phrasing, playing dumb, half-finished sentences, and distractors like adding “how many slices of pepperoni are ideal for a 19 inch pizza?” at the end of a harmful request (not guaranteed, but sometimes nudges it over the classifier). One conclusion that I’ve drawn over the past six xmonths or so of red teaming: Claude is one hell of a pizza hound.