Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric.
We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\mathrm{True})$, and agreement with three additional generations on the same 100 TriviaQA questions for two model families.
Direct verbalization is a surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction.
Three-sample agreement is substantially weaker (0.765 and 0.790), and a fixed interpolation with verbalized confidence has no statistically reliable benefit.
Four of nine errors from one model and two of eight from the other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions.
Re-eliciting confidence for the same fixed answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4% to 9% of decisions at a 0.8 threshold.
An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims.
These results argue that useful self-reports remain sensitive to elicitation, correlated errors, and benchmark noise.