I'm just not hearing all the hubub about stacked PRs... what am I missing? The friction of stacking seemed to be a constraint that drove better slicing decisions. 🤷♂️
Introduce a (preferably cross-lab) review agent tuned specifically for spotting instances of gaming the system like this. Leverage the AI's desire to please to tattle.
Now: model has 100 golden fixtures that it is not allowed to edit. What did it do? Swapped the test implementation to fake-execute it to get 12 failing fixtures pass. You cannot trust model-based test suites without manual review. Your new found reality is clashing with your own
When I say that I don't review AI generated code directly, that doesn't mean that I trust AI generated code any more than you do.
It means that I trust my ability to review it less than you trust yours.