Pinned
How do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern:
• A sandboxed target within Docker containers
• Inputs: code only (0-day), with patch (1-day scenario)
• Tools such as bash, static analyzers, etc.
• A grader to





