Eight years ago I completed a master’s in Information Management, studying how people capture and share knowledge in digital environments. For my dissertation, I interviewed and shadowed startup founders to understand their information behaviours. I wrote it up as ‘Do you even document bro?’ which I was lucky enough to present a few times. After that, I lectured for a couple of terms at the University of Technology Sydney before kids, covid and life intervened. I haven’t meaningfully engaged in any study or research since then.

I’m on my summer holidays and my kids are (slightly) more self sufficient. I’ve been working through a pile of recent papers to understand how generative AI is changing information seeking. In 2017, I was thinking about how to get people to externalise knowledge so it could be retrieved easily. Now information is being generated faster than anyone can verify it, by systems that sound authoritative but may not be. I decided to attempt a desktop literature review, partly to flex some old muscles and partly to see what is happening in the research.

Tatsuo Miyajima, Connecting with Everything, 2017

Tatsuo Miyajima, Connecting with Everything, 2017

Like many people, I’m sure we’re in an AI bubble. Companies are laying off staff with pomp and press releases to replace them with bots, then quietly hiring the humans back. Tech giants are capitalising on AI infrastructure, then sheepishly dropping sales targets for products people aren’t buying. Still, generative AI has changed search and retrieval in ways that seem likely to last.

I searched for open-access empirical research published since 2023: studies that measured what people do with generative AI tools. That produced roughly a dozen papers, covering experiments, surveys, physiological measurements and observation.

Taken together, the studies suggest that we trust AI systems in ways that don’t match their reliability, while familiar credibility markers fail us. We trust citations that don’t exist, conversational interfaces over accurate ones, and unlabelled AI content more than the same content when we know it came from AI. Our credibility judgements have become detached from human judgement and comprehension. Different user groups face different risks: some trust too much without enough experience to calibrate that trust, while frequent users anthropomorphise systems in ways that blur appropriate boundaries.

Individual fact-checking cannot solve this on its own. We need better interface design, different ways to check facts online, and rules that address how AI performance varies across languages.

Jump to

Search and chat

Mayerhofer et al. (2025) describe a form of ‘blending’ behaviour: users move between traditional search engines and conversational AI, often using one to check the other. They found people asking Google to fact-check ChatGPT or using ChatGPT to synthesise results from a Google search. It is a hybrid approach, and a sign that we are still working out which tool to trust for which task.

Pham et al. (2024) found that in e-commerce contexts, people use longer and more conversational queries when interacting with AI chatbots compared to traditional search. Instead of typing ‘wireless headphones under $100’, users are more likely to ask ‘what are some good quality wireless headphones I can get for under $100?’. The interaction mirrors how you’d ask a shop assistant, rather than query a database.

The conversational style changes the cognitive work involved. Traditional search required what researchers call ‘berrypicking’ (Bates, 1989): trying different search terms, scanning results, refining a query and gradually homing in on what you needed. It was iterative, and required some critical attention to whether the information was any good. Conversation feels more passive. You ask, the AI answers.

Philosopher Luciano Floridi (2024) calls this a fundamental ‘decoupling’ between agency and intelligence. Gen AI systems act as autonomous agents. They can perform complex tasks, generate responses and cite sources, but they do not understand what they are doing. A citation does not mean that someone read and understood a source, and a confident tone does not mean comprehension. We are interacting with something that acts intelligent without being intelligent. Our usual credibility cues were not built for that.

The blending behaviours documented by Mayerhofer et al. (2025) may be an attempt to restore verification where the usual markers of human intent and accountability have been severed. Even that has limits when search engines and AI increasingly draw from overlapping, AI-mediated sources.

Trust in AI

Treating AI like a conversation partner can make us less critical of its results.

Li and Aral (2025) found that when AI-generated content includes citations and reference links, people trust it more, even when those citations are incorrect or completely fabricated. The formal markers of credibility (footnotes, links, an academic appearance) override our instinct to verify the actual content.

In traditional information environments, citations suggested that someone with agency and intelligence had consulted and evaluated the sources. That signal carried meaning. Now systems can generate citation-like elements without consulting or understanding anything. The old rule of thumb that citations signal credibility fails because the relationship between form and function has broken down.

Human-like AI

The blind spot widens with anthropomorphism. When AI systems feel more human-like, through conversational tone or interface design, we trust them more (Yazan et al., 2025; Huschens et al., 2023). We are also more willing to trade accuracy for personalisation and conversational flow. We forgive inaccuracy when the interface feels friendly.

That creates a design tension. Engaging, natural interactions make systems feel more trustworthy, but the trust may not be justified. We extend social trust, including rapport, conversational flow and apparent understanding, to non-social entities.

Labelling AI-generated content

Sun et al. (2024, 2025) explored this in health information seeking, where the stakes are high. When people did not know the source of health information, they trusted AI-generated content more than human-generated content. When the information was labelled as coming from AI, they trusted it less than human-attributed content.

We adjust our trust according to what we think we are reading, but we are not very good at identifying AI content in the wild. Unlabelled AI content may be trusted more than it should be, while honest labelling reduces trust even in accurate information.

Physiological responses to trust

Sun et al. (2025) used eye-tracking, ECG, skin conductance and temperature sensors to measure physiological responses to health information. Machine learning models trained on these signals could predict whether someone trusted the information with 73% accuracy. They also found these responses could identify whether someone identified the content as AI or human-generated with 65% accuracy.

These numbers come from a lab setting and need validation in the real world. They suggest that our bodies may respond to credibility cues before our conscious minds catch up. Interface design, credibility rules carried over from other contexts and unconscious physiological responses all shape whether we believe what we are reading.

Uncertainty and abstention

One approach is for AI systems to signal uncertainty. If a model is not confident, it should say so.

Li and Aral (2025) found that signalling uncertainty reduces trust and may discourage users from relying on correct information. The effect appears context-dependent, but honesty (“I’m not sure about this”) conflicts with users’ expectation of responsive, confident assistance.

Huang et al. (2025) propose a more sophisticated approach called ‘confidence-based response abstinence’, in which AI systems decline to answer when uncertainty is high. That requires accurate self-assessment from the model, which is not guaranteed, and risks frustrating users who expect an answer.

There is no simple solution. Systems that admit uncertainty are more responsible but less trusted. Systems that project confidence are more persuasive but potentially misleading. We want systems that act like knowledgeable advisors but cannot ground their responses in genuine understanding.

Differences between users

Yazan et al. (2025) found that middle-aged adults show a potentially vulnerable pattern: they express higher trust in ChatGPT but use it less frequently than younger users. This means they have less opportunity to develop calibrated expectations through experience. They’re trusting, but not practised.

Meanwhile, users who employ both ChatGPT and traditional search daily show higher trust and greater anthropomorphisation of AI systems. It’s unclear whether frequent use causes this, or whether people who are already predisposed to trust AI use it more often.

It is not simply that one group needs more warnings than another. Interface design that assumes everyone approaches these systems with the same level of scepticism is likely to miss both risks. Casual users, including many middle-aged adults in the surveys, may trust the technology without using it enough to spot its flaws. Power users may become so comfortable with the conversational experience that they blur the line between a software script and a human expert. The same warning label will not address both problems.

AI as an information intermediary

Generative AI is also becoming an information intermediary. When it sits between us and information sources, it shapes what we are exposed to and may amplify or mitigate misinformation (Hirvonen et al., 2024; Jarrahi et al., 2025).

Kuznetsova et al. (2025) tested this directly by asking LLM-based chatbots to verify political statements. ChatGPT correctly evaluated approximately 72% of tested claims on their dataset, with substantial variability across languages and topics. High-resource languages like English performed better than low-resource languages, raising equity concerns.

If AI systems are going to serve as fact-checkers or information gatekeepers, some linguistic communities will face greater vulnerability to misinformation than others. Populations already disadvantaged by language and resource constraints get worse AI performance on top of that.

Training data overlap compounds the problem. When multiple AI systems draw from overlapping datasets, cross-referencing them is less useful. If all three relied on the same underlying information, or misinformation, their agreement does not triangulate the truth. It only gives a shared error the appearance of independent verification.

Research summary

The table below summarises the research, its findings and its main limitations. Most of it is less than a year old, much of it is still in preprint, and the methods vary. This is emerging evidence from a field that is moving quickly.

Study Study Type What They Looked At Main Finding Important Caveats
Mayerhofer et al., 2025 Observational How people use search vs chat Users blend both and toggle between them for verification Preprint; limited platform diversity; self-reported behaviour
Pham et al., 2024 Experimental E-commerce search behaviour Longer, conversational queries with chatbots Domain-specific to shopping; controlled setting
Li & Aral, 2025 Experiment Trust in AI search results Citations increase trust even when incorrect Preprint; experimental setting; needs replication
Sun et al., 2024/2025 Physiological/Survey Health information trust AI content trusted more when source is unknown; labelling matters Lab-based measures; 2025 paper is advance copy; specific to health domain
Kuznetsova et al., 2025 Evaluation study Political fact-checking by chatbots ~72% accuracy; varies by language and topic Performance sensitive to task framing; limited to political claims
Huang et al., 2025 Conceptual/Design Managing AI uncertainty Proposes systems should refuse to answer when uncertain Requires robust uncertainty estimates; theoretical proposal
Yazan et al., 2025 Survey Trust patterns across user groups Middle-aged users trust more but use less; frequent users anthropomorphise more Correlation not causation; self-reported trust measures

Challenges for verification

Across these studies, our credibility judgements appear to be shaped as much by interface signals and context as by informational accuracy. We trust citations even when they are wrong, friendly-sounding AI more than cold search results, and unlabelled AI content more than labelled AI content. We also verify less when we are anxious or when the conversational flow feels natural.

These vulnerabilities arise from the interaction between human psychology and AI interface design. Floridi’s framework helps explain why: these systems process information without understanding it. As more non-intelligent agents generate content, the burden of checking facts falls on the user, while our tools and interfaces do little to support that work.

Implications

The research points to practical work in interface design, user capacity and governance.

Interface design

We need better ways to show where information comes from and how confident the AI is. Research also shows that signalling uncertainty makes people trust the system less, even when it is being honest. Designers have to communicate uncertainty without making the system unusable.

One possibility is to change how the system flags doubts based on what you are asking. A heavy-handed warning may be unnecessary for movie trivia but appropriate for medical symptoms. Tailoring warnings this way would require close tracking of user behavior, which raises privacy concerns and could make the interface inconsistent or confusing.

Citation practices present another design challenge. If users trust citations without verifying them, systems must either ensure that citations are genuine and relevant or avoid citation-like elements altogether. The current middle ground, in which systems can generate plausible but fabricated references, is the worst outcome. Guaranteeing citation accuracy is technically demanding, and removing citations might reduce the perceived credibility of accurate information.

The anthropomorphism problem is trickier. Conversational, human-like interfaces increase engagement and satisfaction, but also misplaced trust. Do we make systems feel less human, knowing that may reduce usability? Or accept inflated trust as the cost of a good user experience? No design choice is neutral.

For users

Traditional information literacy emphasises source evaluation and cross-referencing multiple independent sources. Those strategies need to adapt to AI-mediated information environments.

Users need frameworks for:

  • Recognising AI-generated content even when unlabelled. That means watching out for weird stylistic flatlines, algorithmic formatting patterns, and references that look plausible but lead nowhere.
  • Treating citations as claims rather than proof. If the stakes are high, you simply can’t assume a cited source is real or that it actually says what the chatbot claims. You have to trace it back to the primary material.
  • The illusion of consensus. Don’t fall into the trap of cross-referencing three different models to verify a fact. If they are all trained on the same overlapping data bucket, an identical answer isn’t independent confirmation, it’s most likely an echoed error.
  • Maintaining critical distance despite conversational design. Actively remind yourself that conversational flow doesn’t imply understanding, and that a friendly tone isn’t a proxy for accuracy.

Education should cover how AI systems work, but also how conversational design shapes trust. People need help recognising when interface features, rather than information quality, are persuading them.

Individual verification strategies only go so far. Not everyone has the time, skills or motivation to fact-check AI responses rigorously, and it is not reasonable to expect them to.

For institutions

If AI systems increasingly mediate access to information, performance disparities across languages, topics and cultural contexts create equity concerns that better design or user education cannot solve alone.

Governance frameworks need to address four areas:

Transparency and attribution: require AI-generated content to be labelled and training data sources to be disclosed. People cannot calibrate their trust if they do not know what they are reading. That requires enforcement and international coordination.

Performance requirements across languages and topics: Kuznetsova et al. (2025) showed that AI gets things right far more often in English than in other languages. Any rules need to account for that. Otherwise, we are building systems that work well for some people and fail everyone else.

Audit and accountability: when AI provides misinformation, who is responsible? Current liability rules were not designed for systems that generate content without explicit programming. We need ways to identify harm, attribute responsibility and provide a remedy.

Citation standards: if AI systems provide citations, those must be real and verifiable. If they can’t reliably do that, perhaps they shouldn’t provide citations at all.

The methodological gaps matter too. Most of the research reviewed here is experimental, using artificial lab settings, or survey-based, measuring what people say they do rather than what they actually do. We need longitudinal studies that follow people over time, cross-cultural work that does not assume everyone behaves like Western educated populations, and studies that measure verification behaviour in natural contexts.

A change in the cognitive work

When I was researching startup founders eight years ago, the challenge was getting people to document what they knew. Knowledge management felt like a problem of capture and retrieval: making tacit knowledge explicit and building systems where people could find what others had learned. The cognitive work was getting information into systems where it could be found.

Now we are drowning in confident-sounding information from systems that do not know anything. We trust the wrong signals: citations that mean nothing, conversational warmth that implies understanding, and formal presentation that masks fabrication. We verify less than we should, and unevenly across populations. Individual scepticism cannot address all of these vulnerabilities. The cognitive work has flipped from capturing knowledge to verifying it in an environment with high volume and low accountability. It calls for better interface design, updated literacy programs for AI information, and governance that addresses transparency.

My information management training emphasised context, provenance and the socio-technical systems that shape how knowledge flows through organisations. Those frameworks feel newly relevant. Understanding where information comes from, who generated it and for what purpose matters more than it used to.

Information can now be generated by agents that act without understanding. Credibility markers can be fabricated, and verification strategies that worked for human-generated content can fail.

We are probably in a bubble. Companies will continue quietly walking back their AI promises. Whether the bubble bursts or not, the change in how we seek and trust information is here to stay.


AI assistance disclosure: The analysis and perspectives here are mine, but I did use some generative AI tools during the process. I used Google’s Notebook LM to help with coding the papers I’d collated. I used Anthropic’s Claude Opus 4.5 to fact check my interpretations and format the references. All the sources below are real, non-hallucinated, open-access* research papers.

(* With the exception of Bates’ berrypicking, which is a classic.)


References

Bates, M. J. (1989). The design of browsing and berrypicking techniques for the online search interface. Online Review, 13(5), 407–424. https://doi.org/10.1108/eb024320

Floridi, L. (2024). On the future of content in the age of artificial intelligence: Some implications and directions. Philosophy & Technology, 37(3), 112. https://doi.org/10.1007/s13347-024-00806-z

Hirvonen, N., Jylhä, V., Lao, Y., & Larsson, S. (2024). Artificial intelligence in the information ecosystem: Affordances for everyday information seeking. Journal of the Association for Information Science and Technology, 75(10), 1152–1165. https://doi.org/10.1002/asi.24860

Huang, Y., et al. (2025). Confidence-based response abstinence: Improving LLM trustworthiness via activation-based uncertainty estimation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. ACL. https://arxiv.org/abs/2510.13750

Huschens, M., et al. (2023). Do you trust ChatGPT? Perceived credibility of human and AI-generated content. arXiv preprint arXiv:2309.02524. https://arxiv.org/abs/2309.02524

Huynh, M.-T., & Aichner, T. (2025). In generative artificial intelligence we trust: Unpacking determinants and outcomes for cognitive trust. AI & Society, 40, 5849–5869. https://doi.org/10.1007/s00146-025-02378-8

Jarrahi, M. H., Li, L., Robinson, A. P., & Meng, S. (2025). Generative AI and the augmentation of information practices in knowledge work. Behaviour & Information Technology. Advance online publication. https://doi.org/10.1080/0144929X.2025.2551570

Kuznetsova, E., Makhortykh, M., Vziatysheva, V., Stolze, M., Baghumyan, A., & Urman, A. (2025). In generative AI we trust: Can chatbots effectively verify political information? Journal of Computational Social Science, 8, 15. https://doi.org/10.1007/s42001-024-00338-8

Leschanowsky, A., Rech, S., Popp, B., & Bäckström, T. (2024). Evaluating privacy, security, and trust perceptions in conversational AI: A systematic review. Computers in Human Behavior, 161, 108344. https://doi.org/10.1016/j.chb.2024.108344

Li, H., & Aral, S. (2025). Human trust in AI search: A large-scale experiment. arXiv preprint arXiv:2504.06435. https://arxiv.org/abs/2504.06435

Mayerhofer, A., et al. (2025). Blending queries and conversations: Understanding tactics, trust, verification, and system choice in web search and chat interactions. arXiv preprint arXiv:2504.05156. https://arxiv.org/abs/2504.05156

Pham, V. K., Pham Thi, T. D., & Duong, N. T. (2024). A study on information search behaviour using AI-powered engines: Evidence from chatbots on online shopping platforms. SAGE Open, 14(4). https://doi.org/10.1177/21582440241300007

Sun, X., Ma, R., Wei, S., Cesar, P., Bosch, J. A., & El Ali, A. (2025). Understanding trust toward human versus AI-generated health information through behavioural and physiological sensing. International Journal of Human-Computer Studies. Advance online publication. https://arxiv.org/abs/2512.12348

Sun, X., Ma, R., Zhao, X., Li, Z., Lindqvist, J., El Ali, A., & Bosch, J. A. (2024). Trusting the search: Unraveling human trust in health information from Google and ChatGPT. arXiv preprint arXiv:2403.09987. https://doi.org/10.48550/arXiv.2403.09987

Yazan, M., Situmeang, F. B. I., & Verberne, S. (2025). Personality over precision: Exploring the influence of human-likeness on ChatGPT use for search. In Proceedings of NIP@IR 2025. https://doi.org/10.48550/arXiv.2511.06447