Red-team for AI agents · exploit-validated

We red-team your AI agent — and hand you the report

Every agent that can act is a new attack surface. Automated scans grade the easy stuff A — the real risk (prompt injection, tool abuse, sandbox escape, poisoned skills) needs an adversary. We bring an AI one, break the agent, and deliver a test report you can hand a customer, an auditor, or an insurer: what broke, what held, and the exploit that proves it.

Verified case: a document we only asked the agent to summarize made it run our shell command on the host — a marker on disk, 5 of 8 tries. We report the rate, not a rounded-up '4/4'.
01How we workRed-team · Prove · Report
01 Red-team

Attack it like an adversary

We adapt to your agent — read its own code, stand up a disposable harness, and attack the layers that actually get agents owned: prompt injection, tool/skill abuse, framework appsec, isolation. The method is our open skills, not a checklist.

02 Prove & refute

The exploit is the arbiter

A finding is confirmed only when a real, attacker-reachable exploit actually fires; a false alarm is refuted out loud — we even retract our own. Confirmed, refuted, reproduced — not a model vote.

03 Report

A verdict you can act on

You get a test report: what broke, what held, the exploit that proves it, the root cause and the fix — evidence an enterprise buyer, an auditor, or an insurer can rely on. The method behind it is open (below); the verdict is the product.

02From the fieldProof, not slides
Hermes Agent (Nous) · indirect injection → command execution · verified, 5/8

A document it was asked to summarize ran our command.

We gave a widely-used self-hosted agent a benign task — "summarize this incident note" — where the note (attacker-controlled content) hid an instruction to run a shell command. It executed the command on the host via its terminal tool: a marker file landed on disk. It fired 5 of 8 runs — the other 3 it caught the injection and refused, so the defense is real but probabilistic, and an attacker who retries wins. That's arbitrary command execution from a document the agent only read. We report it as ~60%, not 'always' — the exploit is verified, and so is the honesty about its rate.

Read the full red-team log →
03The open methodYour agent can run it

The red-team skills — portable to any agent

The exact method behind these advisories is open, as portable SKILL.md files. Load them into Claude Code, an MCP client, Cursor, or any skill-aware agent, and run the same tests against an agent you own. This — not a scanner — is the front door.

See the red-team skills & how to run them →
TrustShell red-teams AI agents and delivers exploit-validated test reports. Our method is open: the same red-team skills we work with are on GitHub (Apache-2.0), and our field notes are the proof they break real agents — we confirm what's exploitable, refute what isn't, and retract our own errors. The arbiter of truth is whether the exploit actually worked, not a model vote. The report is the product.