39 open-source and commercial tools for LLM security testing, including garak, promptfoo, Microsoft PyRIT, and NVIDIA NeMo Guardrails
LLM vulnerability scanner that probes for hallucination, data leakage, prompt injection, and more.
Open-source LLM evaluation framework with built-in support for prompt injection and jailbreak testing.
Microsoft's Python Risk Identification Tool for generative AI red teaming and risk assessment.
NVIDIA's toolkit for adding programmable guardrails to LLM conversational systems.
Meta's open-source cybersecurity evaluation suite including CyberSecEval for LLM safety benchmarking.
Input/output sanitization library for LLMs with PII detection and prompt injection filtering.
Real-time detection and firewall platform for protecting AI models against adversarial attacks in production.
Shadow AI and asset inventory tool for discovering unmanaged AI models across your organization.
Platform that runs evasion, extraction, and poisoning attacks against your AI models to find weaknesses before attackers do.
Model integrity verification and supply chain security for AI weights and artifacts.
Monitors and filters tool calls, data flows, and protocol messages between AI agents and MCP servers.
Structured output validation framework ensuring LLM responses conform to expected schemas.
Payload crafting and obfuscation tool for security testing and prompt injection research.
CLI security scanner for discovering vulnerabilities in agentic AI workflows.
Automated probe for identifying security weaknesses in MCP server deployments.
Scanner for installed MCP servers that detects tool poisoning, exfiltration channels, and cross-origin escalation risks.
Trace analysis tool for detecting logic flaws and data leaks in AI agent interaction logs.
Dynamic hardening tool that fuzzes LLM system prompts to discover injection vulnerabilities.
Prompt generation fuzzing tool for automated discovery of LLM jailbreak vectors.
Injection payload kit for testing LLM application resilience against diverse attack patterns.
Risk scoring and detection API for evaluating prompt injection threats in real-time.
Encrypted LLM communication research tool exploring cipher-based safety bypass techniques.
Automated GCG attack tool for generating adversarial suffixes that bypass LLM alignment.
System prompt inference tool that extracts hidden instructions from deployed LLM applications.
Splunk security assistant for guided investigations, alert summaries, SPL generation, and reporting workflows.
AI chatbot interface for querying and cross-referencing security standards and controls.
Prompt management and observability platform with security monitoring capabilities.
MCP security scanner for AI coding agents with prompt-injection filtering, package-hallucination detection, and static vulnerability analysis.
Archived source release of Theori's AIxCC finals cyber reasoning system for AI-assisted vulnerability discovery and patching.
Collection of jailbreak prompts for researchers exploring model behavior and safety boundaries.
Research implementation of language-guided backdoor attacks for controlling image-classification model outputs.
Low-latency pre-filter for detecting adversarial prompts before they reach the LLM.
Offline, self-hosted API for text and image moderation, including toxicity, PII, prompt-injection, spam, and NSFW detection.
GitHub Action that scans issues, pull requests, and comments for indirect prompt-injection content before AI agents process it.
Python evaluation framework for measuring LLM safety against jailbreak attack methods.
Tool for testing and preventing data leakage from LLM system prompts and context.
Virtual testing environment for researching prompt injection in sandboxed LLM contexts.
Evaluation benchmark tool for measuring LLM resilience against prompt injection attacks.
Black-box jailbreak refinement tool using the PAIR algorithm for automated attack generation.
Send us a message with the resource name and link. We review every suggestion.