I keep asking every model to do these security reviews, but they never find anything important. I think whoever is doing this needs access to much more powerful models and higher spend. The superficial analysis that consumer AIs are doing simply can't find much.
Consumer AIs only finds stuff after somebody out there run a deeper analysis (which is expensive) and that analysis's results gets shared with the other models.
This is what happened in the coldcard case. It took somebody to spend a lot of money to find the bug and once that was found all the other AIs were able to find too. But before it, you could run the same model over and over again and would not find it.