Back to Blog
Jul 10, 2026
5 min read

T3MP3ST: Evidence-First AI Red Teaming With Autonomous Agents

T3MP3ST: Evidence-First AI Red Teaming With Autonomous Agents

AI-assisted penetration testing is becoming easier to try and harder to evaluate. A tool can now produce a convincing attack narrative, generate payloads, run commands, and produce a report. The harder question is whether the output is useful evidence or just a polished transcript.

T3MP3ST sits in that gap. It is an open-source autonomous red-teaming platform built around AI coding agents, local offensive tools, a browser War Room, scope controls, and evidence capture. The practical way to read it is as an ambitious operator console for authorised AI-driven security research, rather than a finished replacement for a penetration testing team.

For Singapore startups and SMEs, that distinction matters. Customers and procurement teams increasingly ask for security evidence, not broad statements that a tool was run. AI testing can help, but only when scope, proof, review, and remediation are handled clearly.

The Product In One Paragraph

T3MP3ST is a self-hosted command centre for running AI-assisted offensive security work. It can use model APIs, local models, or the coding agent already running on a workstation. The product shape is a mission console: define the target, route the work, let the agent plan and call tools, capture outputs, promote supported findings, and keep retest or learning work visible.

T3MP3ST differs from a conventional scanner. A scanner usually has a defined rule set and a predictable output format. T3MP3ST is closer to an orchestration layer around an agent and an arsenal. Its value depends on how well the operator defines the mission, limits the action space, and reviews the evidence.

The most useful idea is provenance. A model-only observation should stay a hypothesis. A finding becomes more useful when it cites captured tool output, a reproducible artefact, a retest, or a human-supplied proof. That is the discipline any autonomous red-teaming product needs if it is going to be used for real assurance work.

The Comparison Landscape

Autonomous security testing is splitting into several lanes.

XBOW is the commercial autonomous offensive-security option most directly associated with continuous exploit validation. Its public positioning is enterprise-oriented: point it at an application, let it find and prove exploitable attack paths, and receive auditable findings with remediation context. XBOW is the more mature managed-product lane, especially for teams that want a vendor-backed service rather than a self-hosted research workbench.

Strix is developer-facing and application-focused. It positions itself as an open-source AI pentesting tool that runs agents against code and applications, validates issues with proof-of-concept exploits, and integrates into CI/CD. For engineering teams that want pull-request or release-gate testing, Strix is a more direct fit than T3MP3ST.

Shannon is narrower and clearer: an autonomous white-box AI pentester for web applications and APIs. It reads source code, tests the running target, and reports findings it can prove. It is attractive when the team controls the codebase and wants source-aware proof-by-exploitation rather than a broader red-team control plane.

PentAGI is another self-hosted autonomous penetration testing platform, but with a heavier product architecture. It includes a web UI, REST and GraphQL APIs, persistent storage, provider flexibility, and a multi-service deployment model. It is closer to a full autonomous testing platform, while T3MP3ST feels more like an aggressive local operator station wrapped around an existing agent.

garak and promptfoo are in a different category. They are better thought of as LLM application testing and red-team tools. They help test prompts, chatbots, RAG systems, model behaviour, jailbreak resistance, data leakage, and compliance risks. They are often the right starting point for AI product teams, because the scope is narrower and easier to fit into a release process.

The Decision For A Security Team

The choice is less about which product has the most dramatic demo and more about the work you need done.

Use a managed autonomous pentesting product when you want an external platform to test applications continuously and provide evidence your security team can triage. XBOW is the obvious comparison in that category.

Use developer-centred tools when the main goal is to catch exploitable web or API issues before release. Strix and Shannon are easier to explain to engineering teams because they map directly to code, builds, applications, proof-of-concept findings, and remediation.

Use garak or promptfoo when the target is an LLM feature rather than a general application. If the concern is prompt injection, jailbreaks, data leakage, tool misuse, or RAG behaviour, a narrower LLM testing tool usually creates cleaner evidence than a broad autonomous red-team harness.

Use T3MP3ST when you want a self-hosted AI red-team workbench for authorised research, labs, bug bounty preparation, internal rehearsal, or operator-led experiments. It is most interesting for teams that want to inspect and shape the workflow themselves: scope receipts, tool gates, captured evidence, finding promotion, retest tracking, and local agent integration.

Adoption Notes For Teams

The first evaluation should not be against production. Run T3MP3ST in a lab, a staging environment, or an intentionally vulnerable target. Keep it on loopback unless you have added an access-control layer around the API. Treat every networked or active operation as something that needs explicit scope.

The output review should focus on proof. For each finding, ask whether there is a target, an action, a tool output, a reproduction path, and a retest result. If the answer is missing, the finding may still be useful as a lead, but it should not be treated as verified.

The benchmark story should also be read carefully. T3MP3ST publishes artefacts and verification scripts for its own claims, which is better than asking users to trust a landing page. That still does not replace an independent run in your environment. For a serious evaluation, compare tools by the quality of validated findings, false-positive handling, operator control, evidence export, and the effort needed to turn findings into fixes.

For teams evaluating these tools, the practical output is not the fact that an AI tool was used. The useful output is evidence that supports engineering decisions, customer due diligence, security questionnaires, and internal risk review. A short, well-scoped run with clear artefacts is more valuable than a dramatic autonomous transcript nobody can reproduce.

How Palisade Can Help

Palisade helps teams evaluate AI-assisted security testing without losing the fundamentals: authorisation, scope, evidence, exploit validation, remediation, and retest. For tools such as T3MP3ST, that means setting up safe lab runs, reviewing outputs, comparing alternatives, and deciding where agentic testing should sit beside manual penetration testing.

We also support the surrounding work: AI application security reviews, LLM red teaming, MCP and tool-permission hardening, CI release-gate design, cloud and application security testing, and reporting that can stand up to customer due diligence.

To discuss autonomous security testing, AI red teaming, or evidence-backed penetration testing workflows, book a free consultation.