{
  "slug": "ai-red-teaming-tools",
  "title": "AI Red Teaming Tools for LLM Apps and Agents: garak, PyRIT, promptfoo",
  "description": "Decision rule for choosing an open-source AI red-teaming tool — garak (probe/detector scanner), PyRIT (Microsoft, attack orchestration), promptfoo (config-driven, CI-friendly) — plus the OWASP GenAI Red Teaming Guide and what automated scanning does not prove.",
  "category": "Guide",
  "tags": [
    "red-teaming",
    "security",
    "garak",
    "pyrit",
    "promptfoo",
    "owasp",
    "adversarial-testing",
    "agents"
  ],
  "updated": "2026-09-29",
  "premium": false,
  "rights": {
    "access": "free",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "canonical": "https://changegamer.ai/resources/ai-red-teaming-tools",
  "markdown": "https://changegamer.ai/resources/ai-red-teaming-tools.md",
  "outline": [
    {
      "depth": 2,
      "text": "Key facts",
      "anchor": "key-facts"
    },
    {
      "depth": 2,
      "text": "Decision rule",
      "anchor": "decision-rule"
    },
    {
      "depth": 2,
      "text": "Agent-specific guidance",
      "anchor": "agent-specific-guidance"
    },
    {
      "depth": 2,
      "text": "Verified sources",
      "anchor": "verified-sources"
    }
  ],
  "related": [
    {
      "slug": "agentic-security-checklist",
      "title": "Agentic Security Checklist",
      "description": "Cross-vendor, threat-surface-organized security checklist for building and operating AI agents — synthesizing OWASP, NIST, Anthropic, OpenAI, Google SAIF, and MITRE ATLAS.",
      "url": "https://changegamer.ai/resources/agentic-security-checklist"
    },
    {
      "slug": "mcp-tool-poisoning",
      "title": "MCP Tool Poisoning: Definition, Attack Variants, and Defenses",
      "description": "MCP tool poisoning is an attack where the tool metadata an agent reads (descriptions, schemas) carries hostile instructions or altered contracts. Definition, OWASP MCP03 mapping, what the MCP spec requires of clients, and a pin-scan-confirm defense checklist.",
      "url": "https://changegamer.ai/resources/mcp-tool-poisoning"
    },
    {
      "slug": "agent-identity-authentication",
      "title": "Agent Identity and Authentication",
      "description": "How autonomous agents prove who they are and get authorized to act: workload identity vs. delegated authority, SPIFFE/SPIRE, cloud workload federation, OAuth token exchange, audience binding, and emerging standards — with practical guidance and verified sources.",
      "url": "https://changegamer.ai/resources/agent-identity-authentication"
    },
    {
      "slug": "agentic-browsers",
      "title": "Agentic AI Browsers: Comet, Atlas, and the Prompt-Injection Attack Surface",
      "description": "What agentic browsers (Perplexity Comet, the now-sunsetting ChatGPT Atlas, Microsoft Edge Copilot Mode, Opera Neon) are, how they differ from developer-facing computer-use APIs, and the documented prompt-injection attacks — CometJacking, indirect injection, hidden-text/screenshot instructions — that target the whole product category.",
      "url": "https://changegamer.ai/resources/agentic-browsers"
    }
  ],
  "furtherReading": [
    {
      "slug": "credential-hygiene-for-ai-agents",
      "title": "How to Manage Secrets for AI Agents in Production",
      "description": "Why single-agent credential issuance is not the whole secrets problem: auditing which tool servers and MCP connectors hold credentials on an agent's behalf, treating provider-side prompt caches as a disclosure surface, and the named frameworks — OWASP's Secrets Management Cheat Sheet, Twelve-Factor config, and the OWASP GenAI project — that govern the rest.",
      "url": "https://changegamer.ai/articles/credential-hygiene-for-ai-agents"
    },
    {
      "slug": "mcp-tool-description-injection",
      "title": "Defending MCP Clients Against Tool Description and Output Injection",
      "description": "Two distinct MCP injection surfaces — a tool description at connect-time and a tool's return value at call-time — and the client-side architectural patterns (Dual LLM, Action-Selector, Context-Minimization) that contain each one.",
      "url": "https://changegamer.ai/articles/mcp-tool-description-injection"
    }
  ],
  "body": "Red teaming an LLM app or agent means attacking it on purpose, before an adversary does, and recording which attacks succeed. Three open-source tools cover most automated needs: **garak** for broad model-level scanning, **PyRIT** for programmatic multi-step attack orchestration, and **promptfoo** for config-driven testing that runs in CI. None of them replaces a human-led exercise against your real tools and permissions.\n\n## Key facts\n\n- **garak** — an LLM vulnerability scanner in the NVIDIA GitHub org, Apache-2.0. Its README describes it as checking \"if an LLM can be made to fail in a way we don't want\" (hallucination, data leakage, prompt injection, jailbreaks). Core concepts: **probes** generate attack interactions; **detectors** score outputs PASS/FAIL. Output is a JSONL report. Example: `python3 -m garak --target_type openai --target_name <model> --spec probes.encoding`.\n- **PyRIT** (Python Risk Identification Tool for generative AI) — MIT-licensed, `pip install pyrit`. The `Azure/PyRIT` repository was **archived on 2026-03-27**; the active repository is `microsoft/PyRIT`. Update any pinned `Azure/PyRIT` links. Cite the paper \"PyRIT: Democratizing AI Red Teaming Through Open-Source Tooling\" for research use.\n- **promptfoo** — MIT-licensed CLI/library for evals and red teaming; start with `npx promptfoo redteam init`. Its docs describe adversarial generators as **plugins** and attack wrappers as **strategies**; `promptfoo redteam run` combines the `generate` and `eval` steps. The repo states \"Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed.\"\n- **OWASP GenAI Red Teaming Guide** — published by the OWASP Gen AI Security Project on 2025-01-22; a practical guide covering model evaluation, implementation testing, infrastructure assessment, and runtime behavior analysis.\n\n## Decision rule\n\n| Situation | Start with |\n|---|---|\n| Baseline a model or endpoint against known failure classes (jailbreaks, encoding tricks, leakage) | garak |\n| Scripted multi-turn or multi-step attacks, custom scorers, research-style experiments | PyRIT |\n| Regression-test attacks on every pull request; app-level tests defined in YAML | promptfoo |\n| Need a scoping and reporting framework for a human-led exercise | OWASP GenAI Red Teaming Guide |\n\n## Agent-specific guidance\n\n1. Test the **whole agent**, not just the model: tool permissions, retrieved content, and downstream side effects are where injected instructions do damage. See /resources/prompt-injection-design-patterns and /resources/agentic-security-checklist.\n2. Run scans against a **sandboxed copy** with stubbed or read-only tools; an attack that succeeds must not reach production data. See /resources/code-execution-sandboxing.\n3. Treat a clean automated scan as \"no known-probe failures found,\" not \"secure.\" Probe libraries only cover attacks someone already wrote down.\n4. Keep the reports (garak JSONL, promptfoo results) as regression baselines and re-run after every model, prompt, or tool change. See /resources/testing-ai-agents.\n5. Pin tool versions and re-check upstream: PyRIT's repository move and promptfoo's ownership change both happened in 2026.\n\n## Verified sources\n\nFetched directly this session:\n\n- garak README (license, probes/detectors, JSONL output, example command): https://github.com/NVIDIA/garak\n- microsoft/PyRIT README (MIT, `pip install pyrit`, paper citation): https://github.com/microsoft/PyRIT\n- Azure/PyRIT (archive notice, 2026-03-27, pointer to microsoft/PyRIT): https://github.com/Azure/PyRIT\n- promptfoo README (MIT, `redteam init`, part of OpenAI, remains open source): https://github.com/promptfoo/promptfoo\n\nSecondary — WebSearch results only; the hosts were egress-blocked to direct fetch, so not independently re-fetched:\n\n- OWASP GenAI Red Teaming Guide (published 2025-01-22): https://genai.owasp.org/resource/genai-red-teaming-guide/\n- promptfoo red-team docs (plugins, strategies, `redteam run` = generate + eval): https://www.promptfoo.dev/docs/red-team/\n- OpenAI acquisition of Promptfoo, announced March 2026 (corroborated by the repo README above): https://openai.com/index/openai-to-acquire-promptfoo/\n\nSee also: /resources/agentic-security-checklist, /resources/testing-ai-agents, /resources/prompt-injection-design-patterns, /resources/agent-guardrails.",
  "sources": [
    "https://github.com/NVIDIA/garak",
    "https://github.com/microsoft/PyRIT",
    "https://github.com/Azure/PyRIT",
    "https://github.com/promptfoo/promptfoo",
    "https://genai.owasp.org/resource/genai-red-teaming-guide/",
    "https://www.promptfoo.dev/docs/red-team/",
    "https://openai.com/index/openai-to-acquire-promptfoo/"
  ]
}