Continuous red teaming
for AI agents
Red-team your agents, endpoints, and MCP tools automatically. Get reproducible findings and remediation guidance.

Trusted by teams building with AI
Backed by Large Scale Open Source Research
Maintained by the team behind a 100k+ star repository cataloging real world prompt leaks and jailbreaks. Our probe library is grounded in thousands of documented vulnerabilities observed in the wild, not synthetic test cases.
Your Security Agent
Test a system prompt or a live agent endpoint from the dashboard, or wire continuous scans into your pipeline to catch risks before they merge.
Point ZeroLeaks at a system prompt, a live endpoint, or your tool definitions. It runs a full attack in minutes and shows you exactly what got through.
Scan every pull request that changes how your agent behaves. Results land as checks and merge gates, so nothing regresses into production.
Security for teams shipping AI agents
Catch vulnerabilities before launch, then on every change after.

Agents get compromised through their tools, memory, MCP servers, and permissions, not just the system prompt. ZeroLeaks tests all of it.

Get the exact policy or config change to make, then watch ZeroLeaks rerun the attack to confirm the hole is closed.

Attacks only ever reach sinks we own, and destructive actions are simulated, not executed. Point ZeroLeaks at production without the risk.

Every finding comes with the attack that reproduces it, ranked by severity, in a report you can hand straight to engineering or security.
Unlimited scans on every plan
Pay for your team, not your diligence. No “book a demo” wall.
Frequently asked questions
Email us with any other questions.
What can ZeroLeaks test?
ZeroLeaks can test system prompts, instruction sets, deployed agent endpoints, tool definitions, agent skills, and repository changes that affect prompts. You can assess an agent before launch, test its live interface, or add continuous checks to development workflows.
Does ZeroLeaks test live agents or only prompts?
Both. Prompt scans catch weaknesses before deployment, while live agent scans test a configured HTTP endpoint as a real adversary would, evaluating the agent's responses, tool behavior, authorization boundaries, and potential data leakage across attacks that span many turns.
Which attacks does ZeroLeaks cover?
Coverage includes direct and indirect prompt injection, system prompt extraction, tool hijacking, unauthorized actions, sensitive data leakage, social engineering, encoding bypasses, many shot attacks, grooming across many turns, policy manipulation, and context or reasoning exploits.
How does the automated red team work?
Specialized agents plan attacks, generate probes, evaluate responses, and adapt promising techniques based on what the target reveals. Weak branches are pruned while successful paths are expanded, producing validated findings instead of a static checklist of isolated prompts.
What do I receive after a scan?
Each report includes a security score, findings ranked by severity, attack evidence, affected behavior, and actionable remediation guidance. Reports can be reviewed in the dashboard or exported as a PDF for engineering and security stakeholders.
Can ZeroLeaks run in CI/CD?
Yes. The GitHub App can scan pull requests that change AI behavior and report results as checks or comments. You can also trigger scans through the ZeroLeaks CLI or API to enforce security policies in your existing pipeline.
How is sensitive agent data handled?
Submitted prompts and proprietary instructions are processed transiently during active scans and are not retained as full copies after scan handoff. Stored reports contain scan metadata, scores, remediation guidance, and redacted evidence where needed. Your data is not used to train models.
Which models and providers are supported?
Prompt scans support leading models from major providers through the model options available in the dashboard. Live agent testing is model agnostic because ZeroLeaks evaluates your agent through its endpoint, regardless of which model or framework powers it.

