Most healthcare AI benchmarks test static medical knowledge or evaluate tool-using agents on provider-facing tasks. PatientAgentBench generates synthetic patient records and clinical vignettes, then runs multiturn dual-agent conversations scored by an LLM-as-a-jury panel across



