My experience is similar. Even top models like GPT Astra will make your project less efficient and buggy if you let it run loose.
I see people saying they let multiple AI sessions run on their tablet, barely checking generated code and it is very alien to me.
At the current stage, the dev still has to deeply understand the project architecture and give as much context as possible to the AI.
For not so important features (or projects) it is ok to let the AI go wild and fix things later.