ChangeGamer

← All resources

AI Red Teaming Tools for LLM Apps and Agents: garak, PyRIT, promptfoo

Guide · updated 2026-09-29 · Markdown variant

Decision rule for choosing an open-source AI red-teaming tool — garak (probe/detector scanner), PyRIT (Microsoft, attack orchestration), promptfoo (config-driven, CI-friendly) — plus the OWASP GenAI Red Teaming Guide and what automated scanning does not prove.


Red teaming an LLM app or agent means attacking it on purpose, before an adversary does, and recording which attacks succeed. Three open-source tools cover most automated needs: garak for broad model-level scanning, PyRIT for programmatic multi-step attack orchestration, and promptfoo for config-driven testing that runs in CI. None of them replaces a human-led exercise against your real tools and permissions.

Key facts

Decision rule

Situation Start with
Baseline a model or endpoint against known failure classes (jailbreaks, encoding tricks, leakage) garak
Scripted multi-turn or multi-step attacks, custom scorers, research-style experiments PyRIT
Regression-test attacks on every pull request; app-level tests defined in YAML promptfoo
Need a scoping and reporting framework for a human-led exercise OWASP GenAI Red Teaming Guide

Agent-specific guidance

  1. Test the whole agent, not just the model: tool permissions, retrieved content, and downstream side effects are where injected instructions do damage. See /resources/prompt-injection-design-patterns and /resources/agentic-security-checklist.
  2. Run scans against a sandboxed copy with stubbed or read-only tools; an attack that succeeds must not reach production data. See /resources/code-execution-sandboxing.
  3. Treat a clean automated scan as "no known-probe failures found," not "secure." Probe libraries only cover attacks someone already wrote down.
  4. Keep the reports (garak JSONL, promptfoo results) as regression baselines and re-run after every model, prompt, or tool change. See /resources/testing-ai-agents.
  5. Pin tool versions and re-check upstream: PyRIT's repository move and promptfoo's ownership change both happened in 2026.

Verified sources

Fetched directly this session:

Secondary — WebSearch results only; the hosts were egress-blocked to direct fetch, so not independently re-fetched:

See also: /resources/agentic-security-checklist, /resources/testing-ai-agents, /resources/prompt-injection-design-patterns, /resources/agent-guardrails.

#red-teaming #security #garak #pyrit #promptfoo #owasp #adversarial-testing #agents

Category: Guide

Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.

Machine formats: Markdown · JSON · offers at /api/pricing.json · payment at /api/payment.json. Preview the exact corpus format free as NDJSON.