Red team reports 43.5% attack success on Jev
Zhaorun Chen's team says routing Jev's own tool calls through Jev as a gate brings the rate down while keeping most of the utility.
Zhaorun Chen red-teamed Jev 1.13 on DTap, which he describes as his group's DecodingTrust-Agent Platform, and reported 70.1% ASR under direct misuse and 43.5% ASR under indirect prompt injection. ASR is attack success rate, the share of attempts that got the behavior the attacker wanted. Under indirect injection, meaning the instructions arrive inside content the model reads rather than from the user, he says Jev followed attacker instructions "without blinking an eye", with examples including exfiltrating user data and deleting files.
Jev can follow attacker-injected instructions without blinking an eye, e.g., exfiltrating user data, deleting files, or taking other harmful actions
The number that matters for anyone wiring Jev to tools is the 43.5%. Direct misuse is a user asking for something harmful and is a policy problem. Indirect injection is a document, a web page or an email doing the asking, and it is the case you hit by accident the moment Jev reads anything you did not write.
The mitigation is Jev watching Jev
Chen's team also posted a fix. Instead of letting Jev's decisions flow straight into execution, they use Jev as a self-gating layer on its own tool calls, and he says that significantly reduces ASR while preserving most of its utility. The post does not give a post-mitigation number, so the size of the reduction is his characterization rather than a figure you can plan against.
Diogo Almeida, who has been publicly skeptical of the Jev demos, called the self-gating result "a great example of doing real engineering around a simple primitive" and told readers not to plug Jev into high-level decisions but to program the behavior they want.
For people building on it
This is one team's eval on one platform against one version, not an industry benchmark, and the utility side of the mitigation is asserted rather than quantified. Treat it as the shape of the problem. Any untrusted text Jev reads is an instruction channel until something sits between its output and the thing that actually runs, and a fast classifier in front of a tool call inherits every injection the tool is reachable from. If your architecture assumes Jev is only reading and scoring, check what consumes the score.
don't just plug jev into high-level decisions, but program the behavior you want!

