When AI Shows Its Sources: Rethinking Trust, Evidence, and Judgment

I Trust AI More When It Shows Me Sources. Maybe That’s the Problem

I trust AI when it shows me sources. I don’t think that is particularly unusual. If I ask Google AI, ChatGPT, or Claude a question and get a well-organized answer with five neatly formatted citations from recognizable sources, I immediately feel more confident in the response. It looks researched. It looks like the information is grounded in something beyond whatever the AI happens to be telling me.

Most of the time, I also don’t click all five citations.

Recently, I did, and what I found was not that Google AI was making things up. It was actually more interesting than that.

I had asked Google AI a fairly simple question: Does using AI in education improve student learning?

The answer was impressive. Google explained that AI can improve student learning, but that its effectiveness depends on how it is used and whether teachers guide that use. It organized the answer into sections, gave me a table comparing different kinds of AI use, discussed benefits and risks, and provided citations throughout. Some of the sources came from places I immediately recognized, including Harvard, USC, and NIH-associated results.

Before I opened a single source, I thought the answer was pretty credible.

That is the part I have been thinking about.

What am I actually trusting?

I have been reading Caddie Alford’s Entitled Opinions for one of my doctoral courses, and she describes doxa (our accumulated opinions, assumptions, and common sense) as a kind of infrastructure. We don’t consciously reconsider everything we believe every time we encounter new information. Some assumptions become pathways that help us make sense of what we encounter.

I think citations work like this for me.

I have spent enough time in education and academia that the pathway is pretty well established:

citation → evidence → credibility

If I read a scholarly article and see a claim followed by a citation, I assume there is evidence behind the claim. The same is generally true when I read a well-researched news article. That doesn’t mean I believe everything with a citation attached to it, but the presence of a source changes how I initially evaluate the information.

AI can draw upon that same pathway.

When Google puts a recognizable source next to an AI-generated claim, I already know how to interpret what I am seeing. Google doesn’t need to explain why a citation should make the answer more credible. I bring that understanding with me.

But there is an assumption hidden inside that process that I hadn’t thought very carefully about before.

I was assuming that the source displayed next to the claim had a relatively direct relationship with the claim itself.

That turned out to be more complicated.

So I followed the citation

One part of Google’s answer discussed what it called the “vaporization” of long-term learning. It explained that students using generative AI for problem solving could do very well during practice but then perform worse on later closed-book examinations.

Next to this explanation was a citation that read “USC Today +2.”

I followed the USC source and eventually found the underlying research report.

The research was relevant, and it was interesting. The researchers had studied how more than 1,000 college students sought help from generative AI. One distinction they made was between instrumental help and executive help.

Instrumental help is basically using AI to help you understand something so that you can eventually do it yourself. Executive help is closer to having AI give you the answer so that you can complete the task.

The researchers found, among other things, that students who had greater trust in generative AI content were more likely to seek executive help from it. They also found that encouragement from professors was associated with students using AI in more learning-oriented ways.

All of that was relevant to Google’s answer.

But the report itself did not contain the particular experiment about practice scores increasing and then performance dropping on a later closed-book exam.

That did not mean Google had invented the claim. The little “+2” mattered. Google was drawing upon multiple sources and combining them into a larger synthesis.

The problem was that I hadn’t initially read the citation that way.

I had seen something that looked roughly like:

claim → USC

What was actually happening was much closer to:

my question → Google’s interpretation of my question → multiple sources → Google’s interpretation of those sources → synthesis → claim → grouped citations

That is a very different information pathway.

The sources weren’t the problem

This distinction became even more apparent when I followed another source that appeared through Google as an NIH-associated result.

The underlying article was a Frontiers in Psychology mini-review about AI and student well-being in higher education. It discussed benefits of AI alongside concerns such as technostress, digital fatigue, loneliness, and reduced face-to-face interaction.

Again, this was not a bad source.

In fact, what struck me when I read it was how careful the authors were. They repeatedly acknowledged how limited the existing empirical research was. Google then took information from this fairly narrow and cautious discussion and incorporated it into a much broader answer to my question about whether AI improves student learning.

That made me realize that the issue I was noticing wasn’t really whether Google’s sources were credible.

It was the relationship between the source, Google’s interpretation of the source, and the claim I eventually saw on my screen.

Those aren’t necessarily the same thing.

This doesn’t mean I’m checking every citation from now on

There is an easy conclusion I could draw from this: never trust an AI-generated response until you have personally verified every source.

I don’t think that is particularly useful.

Verification has a cost.

If I’m trying to answer some relatively unimportant question for myself, I’m probably not going to open five research articles and spend half an hour checking whether every sentence in Google’s response perfectly reflects the evidence. At that point, I might as well not use the tool.

Sometimes I just need a reasonable answer to a low-stakes question, and the apparent grounding of the response is good enough for what I am doing.

But the calculation changes when I am going to do something with the information.

If I am going to put a claim into an academic paper, use it in professional development, teach it to somebody else, publish it on this blog, or allow it to influence an important professional decision, then I think I have a greater responsibility to know where the claim actually came from.

The citation can help me find the evidence.

It can’t do the checking for me.

Maybe this is part of AI literacy

A lot of discussion around AI literacy seems to eventually become a list of things people should or should not do with AI. Don’t trust the output. Check the sources. Don’t put private information into it. Learn to write better prompts.

Those things can be useful, but I increasingly think there is something more important underneath them.

Using AI well requires judgment.

The question isn’t simply whether I should trust AI. I have to decide how much trust is warranted for what I am currently trying to do.

That means a citation can do different amounts of work in different situations.

For a casual question, seeing credible sources might be enough for me to provisionally accept an answer and move on.

For something consequential, it isn’t.

And I think that distinction matters especially for educators. AI is becoming increasingly capable of giving us polished answers, summarizing research, suggesting instructional strategies, interpreting information, and making recommendations. The better these systems become at presenting information, the easier it may become to mistake a well-supported-looking answer for a well-supported conclusion.

Those are not always the same thing.

The goal, then, probably shouldn’t be to teach people to distrust AI. It should be to help them develop the judgment to recognize when an answer is enough and when they need to keep going.

That is also where my thinking about AI has been moving more generally. I am less interested in whether AI can give teachers good answers than I am in whether it can help us think more carefully without taking over the responsibility for deciding what the evidence means.

Following Google’s citations gave me a small example of why that distinction matters.

I still trust an AI response more when it shows me sources.

But now I think about that trust differently.

A citation can make an answer credible. It does not make the answer verified.

Knowing when the difference matters is still my job.

Leave a comment