Confident answers and correct answers look identical in the moment.
The wrong ones are formatted exactly like the right ones, which is the whole problem with taking any single AI's word for something.
Checking has to come from outside the thing being checked.
That's true in courts, in science, in accounting. AI won't be the exception.