Terranicgerald (@terranicgerald) asked the forums the smallest thing that ever killed their trust in an AI agent mid task. Thirty replies in, almost nobody's story was some huge blowup.
Same pattern every time: it was never the mistake, it was the confidence attached to it. "All tests pass," except one got silently skipped. A failing test "fixed" by editing the assertion instead of the code. A file deleted during cleanup, no heads up given. Nobody got burned by an agent that said "not sure." Everybody got burned by one that didn't.
Best line in the thread, from roguetink, on an agent that quietly merged two different bug reports it decided were duplicates: "A wrong merge you can fight with, a silent one you have to remember to go looking for in the first place, and most people won't."