
Confidence Is the Cheapest Thing AI Writes — and the Least Trusted
We asked 500 US adults who use an AI chatbot at least weekly what makes them trust an answer. The quality AI writing produces most easily finished last, at 8.8%.
September 15, 2026 · 15 min read
By Alex Li, Founder · Contact the author
We asked 500 US adults who use an AI chatbot at least weekly what makes them trust an answer. The quality AI writing produces most easily finished last, at 8.8%. Here is what finished first, and why it lines up with the handful of techniques that measurably improve AI citation.
The short version
- In a survey of 500 weekly AI chatbot users, "it explains its reasoning clearly, step by step" won trust at 49.8% — more than the next three options combined. A confident, authoritative tone finished last at 8.8%.
- Named sources (22%) and concrete numbers (14%) came second and third. Both are also among the three techniques that lifted AI citation in a peer-reviewed 2024 benchmark.
- Asked to name their single biggest concern about AI reliability, 72.0% chose one of the two options describing confident error: "may hallucinate or provide factually incorrect information" (60.6%) and "may sound authoritative while being wrong" (11.4%). A separate open-text question, coded against a stated rule, produced a comparable figure (358 of 500, 71.6%) — though the two are different measurements and we do not treat one as confirming the other.
Confidence is not a trust signal. It is the thing people are actively guarding against.
What we asked, and what we expected
We commissioned a survey of 500 US adults aged 18–65 who use an AI chatbot — ChatGPT, Gemini, Claude, Perplexity, or similar — at least once a week. The questionnaire ran ten questions; this article reports two of them, and the full instrument, every respondent's demographics and every raw answer are public in the report: the survey questions, options, and results.
The one this article turns on asked:
When an AI assistant answers your question, which of these makes you trust its answer most?
- It names the specific sources it used
- It includes specific numbers or statistics
- It explains its reasoning clearly, step by step
- It is written confidently and sounds authoritative
- It matches what I already believed was true
Our hypothesis was that options 1 and 2 would win. The reasoning was simple: a peer-reviewed benchmark had already shown that citing sources and adding statistics are among the strongest levers for getting a page quoted by an AI engine. If those techniques work on machines, we assumed they would mirror how readers decide what to trust.
The data did not agree. Reporting that plainly matters more than a tidy story, because the actual result is more useful than the one we predicted.
Confidence finished last

| What earned trust | Share | Respondents |
|---|---|---|
| It explains its reasoning clearly, step by step | 49.8% | 249 |
| It names the specific sources it used | 22.0% | 110 |
| It includes specific numbers or statistics | 14.0% | 70 |
| It is written confidently and sounds authoritative | 8.8% | 44 |
| It matches what I already believed was true | 5.4% | 27 |
The winner did not squeak past the field. It beat the runner-up by 27.8 percentage points, and it took more support than the next three options combined.
Now separate that table into two groups.
Evidence someone else can check — named sources plus concrete numbers — took 36%.
Presentation — a confident, authoritative tone — took 8.8%.
Readers are not rewarding how an answer sounds. They are rewarding what they can go and verify. And there is a real gap between those two things: 36% versus 8.8% is a factor of four.
One more result worth noting. The option "matches what I already believed was true" finished dead last at 5.4%. Flattery and confirmation do not read as credibility to this audience either — which is a useful counterweight to the assumption that AI content should simply tell readers what they want to hear.
The reasoning gap
Here is where our hypothesis was most wrong, and where the finding gets interesting.
We expected "specific numbers or statistics" to rank first or second. It ranked third, at 14%. The winner was not a piece of evidence at all — it was visible reasoning: an answer that shows how it got to its conclusion.
That is a different thing from being data-rich. A page can be dense with figures and still be a black box: numbers appear, and the reader has no way to see the path that produced them. What 49.8% of respondents selected is the opposite — an answer where the steps are exposed and followable.

Put the two groups side by side and a pattern shows up that is worth naming: 85.8% of respondents chose something that can be checked. Reasoning you can follow (49.8%), a source you can open (22%), a number you can verify (14%). The remaining 14.2% split between an answer that sounded confident (8.8%) and one that agreed with what the reader already believed (5.4%) — neither of which can be checked at all.
What 72% of respondents are actually afraid of
The most lopsided result in this survey came from a closed-ended question, and it is worth being precise about which question produced it.
Respondents were asked to pick their single biggest concern about the reliability of AI-generated content from five pre-written options. 72.0% picked one of the two that describe confident error.
| Biggest concern | Share |
|---|---|
| The AI may hallucinate or provide factually incorrect information | 60.6% (303) |
| The AI may sound authoritative while being wrong | 11.4% (57) |
| The AI may fail to cite credible or up-to-date sources | 10.4% (52) |
| The AI may present biased or one-sided viewpoints | 9.0% (45) |
| The AI may lack context or nuance for complex topics | 8.6% (43) |
Two things about that table. The 11.4% is the only figure here that measures the tone problem directly: it is respondents who read "sounds authoritative while being wrong" and chose it over four substantive alternatives. That is small, and it is a cleaner number than anything we could derive by interpretation.
The other thing is a caution about the 72.0%. The two options that make it up are adjacent but not identical — one is about being wrong, the other about being wrong confidently. Summing them is defensible as long as you say that you summed them, which is what this paragraph is for. A reader who wants a single unsummed number should use 60.6% for "being wrong" and 11.4% for "being confidently wrong."
The survey also carried an open-text version of the same question. Coded for responses that mention confidence or authority together with fabrication or hallucination, 358 of 500 (71.6%) described an AI asserting something false with confidence. That is close to the closed-ended 72.0%, but it is a separate measurement of a separate question, and we are not treating one as confirmation of the other.
Three of the verbatim responses from the open-text question:
"My biggest worry is that the AI sounds so confident and professional that I might blindly trust it even when it is completely making things up."
"My primary concern is that the AI can sound completely confident and authoritative even when it is providing factually incorrect information, which makes it difficult to catch errors without manually verifying everything."
"My main concern is that the AI will confidently state something as a fact when it is actually just making things up, which is dangerous when I need reliable information for my work."
The runner-up options were far behind: missing or unverifiable sources at 10.4%, and one-sided or biased framing at 9.0%.
Read those two results together and the picture is uncomfortable for anyone publishing AI-assisted content. The single feature an AI writing tool produces most effortlessly — smooth, assured, professionally confident prose — is the thing a majority of readers flagged as their single biggest worry — 60.6% naming outright hallucination and a further 11.4% naming the combination of confidence and error. Confidence is not neutral in this data. It is attached to the risk, not to the reassurance.
Where reader trust and AI citation overlap
So far this is a story about readers. Your content also has a second audience that never reads it: the retrieval systems that decide whether to quote you at all.
Those two audiences are measured in completely different ways, so we should be careful about claiming they think alike. But there is an overlap worth noting, because it is narrow.
The supply side has been benchmarked. In GEO: Generative Engine Optimization — published at KDD 2024 and available as arXiv 2311.09735 — the authors evaluated nine content modifications across 10,000 queries. Their conclusion, in their own words:
"Our top-performing methods, Cite Sources, Quotation Addition, and Statistics Addition, achieved a relative improvement of 30-40% on the Position-Adjusted Word Count metric and 15-30% on the Subjective Impression metric."
Their best-performing methods improved on the baseline by 41% and 28% on those two metrics.

Line up the two studies:
| Rewards | Ignores | |
|---|---|---|
| AI engines (KDD 2024 benchmark) | Named sources, quotations, concrete statistics | Keyword stuffing — which "offer[s] little to no improvement" |
| Readers (our 500-person survey) | Visible reasoning, named sources, concrete numbers | A confident, authoritative tone (8.8%) |
The overlap is named sources and concrete numbers. That is the only pair of levers that both studies reward. Reasoning that a reader can follow has no equivalent on the supply side — a retrieval system does not care whether your argument is easy to trace. And each side has something it actively punishes: the benchmark found keyword stuffing offered little to no improvement, and measured it 10% below the baseline on Perplexity.ai; our survey found an authoritative tone chosen by just 8.8% of readers, ahead of only one other option.
If you only have room to optimize for one thing, optimize for the overlap: sources a reader can open, and numbers a reader can check.
Why "structured for AI" is only the entry ticket
There is a popular idea that getting cited by AI is a formatting problem — clean headings, answer-first paragraphs, a tidy FAQ block. That advice is not wrong, but it describes the entry ticket, not the competition.
Our own AI citeability checker measures exactly this layer, and the tool's own description is deliberate about its limits: it checks whether content is structured so an AI can understand it and would likely cite it — "content-level, not actual citation."
That distinction is the whole point. Structure is what makes your page legible to a retrieval system. It does not give that system a reason to prefer your page over the twelve others with the same structure. The reason has to come from something only you have — and the two studies above agree on what that is.
How to make content verifiable instead of merely confident
Three moves, in order of how much they change the outcome.
1. Attach every important number to a named source the reader can open. Not "studies show" and not "industry data suggests." Name the study, name the year, link it. A benchmark number with a citation behind it is worth several uncited numbers, and it is the single technique that appears at the top of both studies.
2. Run your own research, and publish it as a report — not as a claim. This is the highest-leverage version of move one, because it produces a source that did not exist before and that nobody else can cite instead of you. It is also the only way to get the one thing a competitor cannot copy. Our own survey in this article is the example: ten questions, 500 respondents, two reported, one permanent source.

3. Show the reasoning, not just the conclusion. This was the winner at 49.8%, and it is the move most content teams skip. In practice it means writing the "because" — the assumption you tested, the option you rejected and why, the number that surprised you. It is also the reason this article reports that our hypothesis failed. A page that only shows conclusions asks to be believed. A page that shows its path asks to be checked, and readers overwhelmingly chose checking.
Note what is not on this list: rewording for a more authoritative tone, adding "expert" phrasing, or smoothing the prose. Those are the cheapest things to produce and, at 8.8%, the least rewarded.
Frequently asked questions
Is a confident tone bad for content? Not harmful by itself, but it earns almost nothing. In our survey of 500 weekly AI chatbot users, "written confidently and sounds authoritative" was chosen by 8.8% as the quality that makes them trust an answer — last among five options, behind named sources (22%) and concrete numbers (14%). A confident tone is fine as a delivery mechanism. It is not evidence, and it does not read as evidence.
What makes an AI answer trustworthy to readers? Visible reasoning first, at 49.8% — ahead of named sources (22%), concrete numbers (14%), confident tone (8.8%), and agreement with the reader's existing beliefs (5.4%). Combined, 85.8% of respondents picked something that can be independently checked.
Do people actually distrust AI-generated content? They distrust unverifiable confidence in it. Asked to pick their single biggest concern about AI reliability from five options, 72.0% of our 500 respondents chose one of the two describing confident error — "may hallucinate or provide factually incorrect information" (60.6%) or "may sound authoritative while being wrong" (11.4%). Those were two separate options; the 72.0% is their sum. Concerns about missing sources (10.4%) and bias (9.0%) came a distant second and third.
Why does ChatGPT cite some sources and not others? At the level of levers we can measure and test, three content modifications consistently improved AI visibility in the KDD 2024 Generative Engine Optimization benchmark: citing sources, adding quotations, and adding statistics. That same benchmark evaluated nine modifications in total, and found keyword stuffing offered little to no improvement — 10% below the baseline on Perplexity.ai. Engine retrieval behavior changes frequently, so treat a 2024 benchmark as a directional result about what makes a source quotable rather than a fixed specification.
Do statistics make content more citable? The benchmark evidence says yes, and so does reader behavior — but less than most people assume. Concrete numbers ranked third with readers at 14%, behind visible reasoning at 49.8%. Statistics are necessary and not sufficient: they need attribution, and they need to be visible to someone following your argument rather than dropped into a paragraph.
Can structure alone get my content cited? No. Structure makes your page legible to a retrieval system, which is why we built a citeability checker around it — but legibility is an entry ticket, not a tiebreaker. When twelve pages share the same clean structure, the system still needs a reason to pick one. That reason comes from evidence only your page contains.
Do I need a research budget to publish citable content? No, but you do need one original observation. Original data is the strongest form because it creates a source that did not previously exist and that no competitor can cite instead of you. A short survey of a few hundred respondents in your buyers' segment is enough to produce a permanent, citable asset. The survey behind this article ran ten questions to 500 respondents and we report two of them — the outcome ranking and the biggest-concern question with its open-text follow-up.
What is the most common mistake when publishing AI-assisted content? Publishing prose that sounds certain while containing nothing checkable. It is the default output of every AI writing tool, and it is precisely what the 72.0% chose — with 11.4% naming the confident-but-wrong combination outright, ahead of only bias and lack of nuance. The fix is not better prompting for tone. It is adding the two things that both readers and retrieval systems reward: sources a reader can open, and numbers a reader can verify.
Check your own content before you publish it
Structure is the part you can check today. Run a page through our free AI citeability checker to see whether an AI can parse it cleanly — then treat a passing score as the floor, not the goal.
The part that decides whether you get quoted is the part the checker can't see: whether your page contains evidence that exists nowhere else. That is what Chatgpt Grow is built to produce — a customer question, original research behind the answer, and one published article a day, on autopilot.
Method note. The reader survey was fielded 15 September 2026 with 500 US adults aged 18–65 who use an AI chatbot at least weekly, drawn from a candidate pool of 96,162. Sampling was quota-matched across age band, gender, education level, personal income band, US census region, and occupation group (23 categories under the US census classification). The reported run contains 500 completed responses, zero recorded failures, and no data-quality warnings.
Three disclosures. The survey platform is our own product (Moevox), so we commissioned this research on the tool we sell; the full instrument, every respondent's demographics, and every raw answer are published at the report link below so the result can be checked rather than taken on trust. The questionnaire contained ten questions and this article reports two of them — the trust-driver ranking and the biggest-concern question with its open-text follow-up; the other eight are context for the report and are not used in any claim here. The open-text coding rule is stated in the text above rather than left to the reader to infer, because a share derived from coding is only as good as the rule behind it. Full results: moevox.com report. The citation benchmark is Aggarwal et al., "GEO: Generative Engine Optimization," KDD 2024 (arXiv 2311.09735). The two studies measured different things — reader trust and AI citation — and are presented as converging evidence, not as the same measurement.
Chatgpt Grow