Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven EvaluationPrescriptive constitutional definitions and AI-driven evaluation for consistent golden labels in content moderation pipelines.May 2026arXiv →
Cisco Integrated AI Security and Safety Framework ReportUnified framework spanning content safety failures, model-level attacks, and networked systems with embedded agents.Dec 2025arXiv →
Toward Quantitative Modeling of Cybersecurity Risks Due to AI MisuseNine cyber risk models analyzing AI uplift to offensive operations as a function of benchmark performance.Dec 2025arXiv →
Death by a Thousand Prompts: Open Model Vulnerability AnalysisSecurity assessment of open-weight LLMs revealing 2-10× higher attack success in multi-turn scenarios.Nov 2025arXiv →
A Framework for Rapidly Developing and Deploying Protection Against LLM AttacksProduction-grade defense system integrating threat intelligence, data platforms, and rapid deployment for evolving LLM threats.Sep 2025arXiv →
LLM Cyber Evaluations Don't Capture Real-World RiskPosition paper proposing a risk assessment framework that incorporates threat actor behavior and impact potential.Feb 2025arXiv →
Recent Projects
VigilDetection system for prompt injections, jailbreaks, and other risky LLM inputs. Layered defense approach.PythonGitHub →
CascadeFacilitates conversations between two LLMs with optional human-in-the-loop for alignment research.PythonGitHub →
QubitMinimalist blogging platform. Simple, fast, focused on writing.PythonGitHub →