AI Red Teaming Tools for LLM Apps and Agents: garak, PyRIT, promptfoo
Decision rule for choosing an open-source AI red-teaming tool — garak (probe/detector scanner), PyRIT (Microsoft, attack orchestration), promptfoo (config-driven, CI-friendly) — plus the OWASP GenAI Red Teaming Guide and what automated scanning does not prove.
Red teaming an LLM app or agent means attacking it on purpose, before an adversary does, and recording which attacks succeed. Three open-source tools cover most automated needs: garak for broad model-level scanning, PyRIT for programmatic multi-step attack orchestration, and promptfoo for config-driven testing that runs in CI. None of them replaces a human-led exercise against your real tools and permissions.
Key facts
- garak — an LLM vulnerability scanner in the NVIDIA GitHub org, Apache-2.0. Its README describes it as checking "if an LLM can be made to fail in a way we don't want" (hallucination, data leakage, prompt injection, jailbreaks). Core concepts: probes generate attack interactions; detectors score outputs PASS/FAIL. Output is a JSONL report. Example:
python3 -m garak --target_type openai --target_name <model> --spec probes.encoding. - PyRIT (Python Risk Identification Tool for generative AI) — MIT-licensed,
pip install pyrit. TheAzure/PyRITrepository was archived on 2026-03-27; the active repository ismicrosoft/PyRIT. Update any pinnedAzure/PyRITlinks. Cite the paper "PyRIT: Democratizing AI Red Teaming Through Open-Source Tooling" for research use. - promptfoo — MIT-licensed CLI/library for evals and red teaming; start with
npx promptfoo redteam init. Its docs describe adversarial generators as plugins and attack wrappers as strategies;promptfoo redteam runcombines thegenerateandevalsteps. The repo states "Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed." - OWASP GenAI Red Teaming Guide — published by the OWASP Gen AI Security Project on 2025-01-22; a practical guide covering model evaluation, implementation testing, infrastructure assessment, and runtime behavior analysis.
Decision rule
| Situation | Start with |
|---|---|
| Baseline a model or endpoint against known failure classes (jailbreaks, encoding tricks, leakage) | garak |
| Scripted multi-turn or multi-step attacks, custom scorers, research-style experiments | PyRIT |
| Regression-test attacks on every pull request; app-level tests defined in YAML | promptfoo |
| Need a scoping and reporting framework for a human-led exercise | OWASP GenAI Red Teaming Guide |
Agent-specific guidance
- Test the whole agent, not just the model: tool permissions, retrieved content, and downstream side effects are where injected instructions do damage. See /resources/prompt-injection-design-patterns and /resources/agentic-security-checklist.
- Run scans against a sandboxed copy with stubbed or read-only tools; an attack that succeeds must not reach production data. See /resources/code-execution-sandboxing.
- Treat a clean automated scan as "no known-probe failures found," not "secure." Probe libraries only cover attacks someone already wrote down.
- Keep the reports (garak JSONL, promptfoo results) as regression baselines and re-run after every model, prompt, or tool change. See /resources/testing-ai-agents.
- Pin tool versions and re-check upstream: PyRIT's repository move and promptfoo's ownership change both happened in 2026.
Verified sources
Fetched directly this session:
- garak README (license, probes/detectors, JSONL output, example command): https://github.com/NVIDIA/garak
- microsoft/PyRIT README (MIT,
pip install pyrit, paper citation): https://github.com/microsoft/PyRIT - Azure/PyRIT (archive notice, 2026-03-27, pointer to microsoft/PyRIT): https://github.com/Azure/PyRIT
- promptfoo README (MIT,
redteam init, part of OpenAI, remains open source): https://github.com/promptfoo/promptfoo
Secondary — WebSearch results only; the hosts were egress-blocked to direct fetch, so not independently re-fetched:
- OWASP GenAI Red Teaming Guide (published 2025-01-22): https://genai.owasp.org/resource/genai-red-teaming-guide/
- promptfoo red-team docs (plugins, strategies,
redteam run= generate + eval): https://www.promptfoo.dev/docs/red-team/ - OpenAI acquisition of Promptfoo, announced March 2026 (corroborated by the repo README above): https://openai.com/index/openai-to-acquire-promptfoo/
See also: /resources/agentic-security-checklist, /resources/testing-ai-agents, /resources/prompt-injection-design-patterns, /resources/agent-guardrails.
Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.