MENU
Blog

How do AI engines choose what to cite? We tracked 2,429 answers to find out

AI engines do not cite the biggest brand, the highest domain authority, or the most popular page. They cite whoever published the clearest answer to the question being asked, and most of the time, nobody credible has. We know because we spent three months logging where four AI systems get their answers about one contentious industry: 2,429 answers, roughly 19,100 cited links.

The industry was data centers. We chose it because it is one of the most contested subjects in American local politics right now, the questions people ask about it are urgent and specific, and almost every organization with a stake in the answers has published almost nothing that an AI engine can use. That combination makes it a nearly perfect laboratory for watching how AI answers actually get made. If you write content for a living, what we found should change how you think about your job.

What did we actually track?

From May 4 to July 13, 2026, we used Meltwater GenAI Lens to put 20 recurring questions about data centers to ChatGPT, Google AI Mode, Google AI Overviews, and Perplexity, on roughly 35 collection dates. The questions covered environment, energy, water, community impact, jobs, health, and policy. Every answer was logged with its sentiment, the organizations it named, and every source it cited. That produced 2,429 answers carrying about 19,100 cited links, which is a large enough pool to see patterns that single spot-checks miss.

Who does AI actually cite?

Facebook, YouTube, and Reddit are the three most-cited domains in the entire dataset. Ahead of the Department of Energy. Ahead of the International Energy Agency. Ahead of every news outlet, which collectively account for only about 6% of cited links. Social platforms together supply roughly one in seven of all citations.

And these are not neutral links. The Facebook citations include posts from community groups organized against data center construction, including one literally named “Say No to Data Centers.” On the questions a worried resident actually asks, like what it is like to live near a facility, whether it is safe, and what happens to noise and property values, community opposition content is a primary source. An advocacy page titled “The Dangers of Data Centers” is the single most-cited source on the health question, with 65 citations on that prompt alone.

Read that again from a content writer’s perspective: the engines wanted to answer these questions, found almost nothing from institutions or industry, and cited whoever had bothered to write anything at all.

Why don’t the biggest companies get cited?

The companies at the center of the story are its protagonists, not its sources. The hyperscalers that build and run these facilities are named in well over a thousand answers combined, yet across 19,100 cited links their own domains appear only a few dozen times, and several of them never appear at all. Major data center operators barely register either. The organizations with the most at stake, and the largest content budgets on earth, have effectively no voice in the answers.

Being talked about and being cited are two different games. Mentions make you the subject of the story. Citations let you shape it. Almost nobody in this category is winning both.

So who beats them? Companies you have never heard of

A single blog post from STAX Engineering, an emissions-capture equipment vendor, drew 131 citations over the quarter, making STAX a top-cited authority on data center emissions. EziBlank, an Australian manufacturer of server-rack blanking panels, owns the question “is it safe to live near a data center” through a handful of FAQ-style pages whose URLs are literally the question. Stream Data Centers’ state-by-state tax-incentive glossary is cited in 43% of all answers on the tax question. Hanwha’s data center arm owns the energy-sourcing question.

None of these companies are household names. All of them out-cite the hyperscalers by an order of magnitude, on the strength of pages that simply answer a specific question directly. The bar for owning an AI-answered question is strikingly low, because almost nobody is trying to clear it.

The same dynamic rewards institutions that publish clean numbers. Lawrence Berkeley National Laboratory’s data center electricity figures appear in 150 to 240 separate answers each. Nobody outranked them. They published the number first, with methodology, and now they are simply the answer.

Does every AI tell the same story?

No, and the differences are large enough to change what you should publish. ChatGPT behaves like a cautious institutional analyst: it cites the IEA, the Department of Energy, and the EPA, uses almost no social content, and hedges its conclusions. The Google surfaces behave differently: AI Mode and AI Overviews make assertive claims sourced heavily from social and advocacy content, and they are far more negative on contested questions. Since those two sit inside Google Search, the most negative framings also reach the largest audience by far. Perplexity sits in between and moves fast: its tone on this topic flipped sharply in a single month.

One number that should recalibrate how you think about authority: across the whole dataset, the same underlying facts appear everywhere, but which sources get quoted, and what tone the answer takes, depends mostly on which engine you ask.

What happens when nobody answers? The vacuum fills itself

Over just ten weeks, negative answers rose from 15.6% of the dataset in May to 24.2% in early July. Facebook’s share of all cited links tripled over the same period. Answers about electricity prices went from 39% to 62% negative; answers about local water supplies went from 21% to 53%.

Nothing on the industry side entered the citation pool that quarter to push back. That is the mechanism in one sentence: as one side keeps publishing and the other stays silent, retrieval-based engines index more of what exists, and the answers harden. Silence does not produce neutrality. It produces someone else’s answer, delivered at scale, inside the default search experience.

What should content writers take from this?

  • Answer real questions, literally. The pages that win are shaped like the question people ask, down to the URL. Topic labels lose to questions.
  • Domain authority is not the lever. A rack-panel manufacturer beats trillion-dollar companies because it published the answer and they did not.
  • Mentions and citations are different metrics. Track them separately, because you can dominate one and be invisible in the other.
  • Pick your engine deliberately. Institutional, hedged content reaches ChatGPT. Direct, extractable, fresh content wins the Google surfaces, where the audience is.
  • Vacuums do not stay empty. If your organization will not answer the hard questions about your category, the engines will find someone who will, and you will not like their framing.

Get the full playbook

Video is one of eight channels in the AEO Field Guide, the complete playbook for getting content cited by AI: seven universal rules, a 60-second pre-publish checklist for every channel, and the research behind all of it. Free.

Get the free field guide

Frequently asked questions

No. Popularity barely matters. In our data, the most-cited sources on several questions were small vendors and advocacy pages with tiny audiences, and outside research finds view counts correlate with citations at effectively zero.

Sometimes one page. STAX Engineering's 131 citations came from a single blog post. Owning a question durably usually takes a library of focused pages, but the entry price is far lower than most brands assume.

The mechanism is not. Every category we have measured shows the same pattern: engines cite the clearest available answer, and most categories have more unanswered questions than answered ones. The data center category is simply an extreme case of the vacuum.

About this research

Hahn tracked 2,429 AI answers to 20 recurring questions about data centers across ChatGPT, Google AI Mode, Google AI Overviews, and Perplexity, May 4 to July 13, 2026, using Meltwater GenAI Lens, logging roughly 19,100 cited links. This tracking ran as a companion study to the Escalent + Hahn Data Center Communications Playbook, a 13-state survey of 3,417 U.S. consumers on public attitudes toward data centers.

Both the research and this article were produced in collaboration with AI, and I think that is worth saying plainly. AI helped me analyze a dataset of thousands of answers and tens of thousands of cited links, and helped me draft this piece. That collaboration is the only practical way one person does work at this scale. The judgment calls, the verification of every figure against the underlying data, and the conclusions are mine.