Pinned
New benchmark: RADAR evaluates 21 frontier models on vulnerability identification across 41 real-world XBOW codebases, scored against human expert ground truth.
Primary metric is recall (asymmetric cost structure in security). Top result: 62.4%.
The recall/F1 ranking inversion


