I've spent the last six months writing skills for AI coding agents. Somewhere around month three, I started noticing something uncomfortable: some of our best-intentioned guidance was making things worse. Not on every task — just on enough tasks to make me question whether we were actually helping.
I couldn't prove it. I had gut feelings and a few suspicious eval runs, but nothing concrete enough to point at and say "this is why."
Then I read this paper.