skipToContent
Back to all posts

Sources Behind AI Answers: Where Mid-Sized Companies Can Actually Get In

August 14, 2026
By Michael Kaiser
GEOAI VisibilitySourcesWikipediaSME
Several nodes of differing size whose light streams converge into a central answer card, two of the largest ones sitting behind a barrier

The advice sounds plausible and turns up in almost every consultation on the subject: you need a Wikipedia entry, then ChatGPT will cite you too. For the vast majority of mid-sized companies this advice is useless, for a reason you can verify in two minutes: they do not meet the inclusion criteria and never will.

Key takeaway: AI systems answer predominantly from other people's sources, not from your website. An analysis of more than 680 million citations shows that around 15 domains supply roughly 68 percent of all citations, led by Reddit, Wikipedia, YouTube and LinkedIn. The catch: a mid-sized company can hardly get into the top two. The German Wikipedia requires 1,000 employees, 100 million euros in revenue or 20 business locations. What works instead is less spectacular and slower, but it is achievable: consistent information across several independent sources.

Which sources do AI systems actually answer from?

From a few large ones and a great many small ones. The AI Platform Citation Source Index 2026 consolidated more than 680 million individual citations from six studies conducted between August 2024 and April 2026. The result: around 15 domains supply roughly 68 percent of all citations produced by ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews. Reddit, Wikipedia, YouTube and LinkedIn lead the field.

The distribution differs markedly by system. That is the more useful finding in practice, because it determines where effort is worth spending at all.

SystemNotably strong sourceShare of top sources
ChatGPTWikipediaaround 48 %
PerplexityRedditaround 47 %
Google AI OverviewsReddit, just ahead of YouTubearound 21 %
GeminiReddit barely matters, Medium strongeraround 3 %

So when someone says Reddit is decisive, the question to ask back is: decisive for which system? For Perplexity it holds. For Gemini it is close to meaningless.

Why the Reddit dominance in these studies is overstated

Because two different things are measured and given the same name. The spectacular percentages almost always refer to the share within the ten most frequent sources. Look at all citations instead and the picture changes completely: an Evertune analysis covering 200 million prompts over five months finds that even a platform's most-cited domain rarely accounts for more than 5 percent of all citations.

Both figures are correct. They simply answer different questions. One asks who sits at the top. The other asks how much of the whole that actually is. And the second answer matters more in practice, because it says the remaining 95 percent are spread across a very long tail of smaller sources. That tail is the reachable part.

Skip this distinction and you arrive at the wrong strategy: all effort on two platforms you cannot get into anyway, instead of on the many small ones that together supply the larger share.

Will my company make it into Wikipedia at all?

Probably not, and that can be settled in two minutes. The German-language Wikipedia has written inclusion criteria for companies. Either at least one is met, or none is.

CriterionThreshold
Headcountat least 1,000 full-time positions
Annual revenueat least 100 million euros
Stock listingon a regulated market or comparable segment
Business locationsat least 20 sites, branches or outlets

There is a fifth route, and it is the only one open to smaller companies: a market-dominating position or an innovative pioneering role. Both, however, must be evidenced by external sources, meaning quality media or academic publications, not by your own press release. Anyone holding that evidence usually no longer needs Wikipedia urgently, because the coverage itself is already the source models draw on.

For a typical mid-sized company with 30 staff at one location this means the route is closed. That is not a question of effort but of the rule.

What happens if I try anyway?

At best nothing; at worse a new kind of damage that was not there before.

First the formal side: anyone paid for a contribution or writing it on instruction must disclose that under the Wikimedia Foundation's terms of use, naming employer, client and affiliation. This applies to employees writing about their own employer too. Violations regularly lead to the account being blocked. An agency offering to handle the entry discreetly is offering a breach of the rules.

The second point weighs heavier. If an article is created about a company that does not qualify, a deletion request generally follows. The deletion discussion is public, it remains permanently accessible, and it is indexable. What stands at the end is then not the desired entry but a publicly archived discussion establishing that the company fails the relevance criteria. As a source for a language model, that is worse than no entry at all.

Which directories actually count?

Those with their own editorial oversight or a genuine barrier to entry. Directories anyone can join in five minutes for free carry little weight with models, because they constitute no independent confirmation. An entry is useful where someone checks submissions, where reviews come from verified customers, or where membership is a precondition.

Type of sourceEffortEffect
Google Business ProfileLowFoundation for local queries, widely evaluated
Chambers and associationsLowIndependent confirmation, often underrated
Review platforms with verified customersMediumWeighted highly because not self-controlled
Software and vendor directories with editorsMediumOnly relevant if the category fits the offering
Trade media and industry pressHighStrongest effect, because editorial and citable
Free listing directories without reviewLowPractically ineffective, harmful amid data chaos

The last point warrants a warning. Anyone entering their company in thirty aggregator directories, using the old address in some and the new one in others, creates exactly the problem they set out to solve: contradictory information across the web. How that turns into false statements in AI answers is covered in our article on false AI claims about your company.

What works instead for mid-sized companies?

Consistency across several independent sources. It is unspectacular, it cannot be done in a week, and it is the only lever that actually works for a company without Wikipedia relevance.

The mechanism behind it: models treat a statement as established when several mutually independent sources say the same thing. A self-declaration on your own website is one source. The same statement additionally in a chamber register, a trade article and a review profile is something else: a repeatedly confirmed fact.

That yields a sequence, and it can be worked through in that order.

  1. Unify your own information. Legal name, founding year, location, services, contacts. Identical everywhere, including on old subpages and in PDFs. As long as two versions are findable, the model decides for itself which one is true.
  2. Make your website machine-readable. Schema.org markup for the organisation, people and articles, so the information need not be inferred from prose.
  3. Clean up existing third-party sources. Chamber entry, Google profile, review profiles, old press articles. There are usually more than you remember, and usually at least one is out of date.
  4. Create new independent sources. Trade articles in your industry's publications, talks with documented announcements, association work. Slow, but durable.
  5. Only then consider platforms. And only where your audience actually is. A trades business has no business on Reddit; a B2B software vendor might.

The order rests on a simple consideration: the first three steps cost little and prevent damage. The last two cost a lot and build something. Start at step five and you merely amplify the contradictions left unresolved in steps one to three.

How long until source work takes effect?

Longer than it can be sold as. With systems that search the web live, answers can change within weeks once a corrected or new source is indexed and readily findable. With answers from the training state, nothing changes before the next training cycle, and that is months away.

A realistic observation period is therefore a quarter before it is even possible to judge whether a measure achieved anything. And judging means measuring, not asking once: answers vary so much between individual queries that a single measurement says nothing. How to build a defensible measurement is covered in our article on measuring AI visibility.

Anyone promising faster results is promising something they cannot deliver. A mention cannot be bought, only made more likely.

Frequently asked questions

Do I need a Wikipedia entry to appear in AI answers? No, and most companies will not get one anyway. The German Wikipedia requires 1,000 employees, 100 million euros in revenue, a listing on a regulated market or 20 business locations. Wikipedia is the strongest single source for ChatGPT, but across all citations it accounts for only a small share. The rest is spread across many reachable sources.

Can I write a Wikipedia article about my own company? Technically yes, sensibly only if the relevance criteria are met. Anyone paid for it or writing on instruction must disclose that under the terms of use; violations lead to the account being blocked. Without relevance, the attempt usually ends in a publicly archived deletion discussion that is itself findable.

Is it worth listing in as many directories as possible? No, and in case of doubt it does harm. What matters is not the number of entries but that they all say the same thing. Directories without editorial oversight or an entry barrier carry little weight, because they constitute no independent confirmation. One contradictory entry with an old address does more damage than ten additional entries do good.

Why is Reddit named the most important source in nearly every article? Because the percentages quoted usually mean the share within the ten most frequent sources, not the share of all citations. Across all citations, even the strongest domain rarely exceeds 5 percent. On top of that, its importance varies sharply by system: dominant for Perplexity, practically meaningless for Gemini.

What is the single most effective measure? Eliminating contradictions in your own information before building anything new. As long as two different versions of your legal name, address or services are findable, the model decides for itself which is correct, and it does not always decide in your favour.

How do I tell whether a source matters for my company at all? Two questions: does anyone review what gets published there, and is my audience actually present? If both are yes, the effort pays. If either is no, the entry is at best ineffective.

Conclusion

The sources AI systems answer from are largely not your own website. That is the uncomfortable starting point. The common conclusion drawn from it, that you simply need to be present on Wikipedia and Reddit, does not help mid-sized companies: one route is blocked by inclusion criteria, the other leads away from your actual audience.

What remains is less spectacular and works more reliably. Put your own information in order, clean up the third-party sources that exist, then patiently create new independent evidence. No tool meaningfully accelerates this, and no agency can guarantee a mention.

Getting started costs a morning: search your company name and list every result that is not yours. Check each one for whether the details still hold. In most cases the list is longer than expected, and at least one entry dates from a time when the address was a different one.

This article reflects the state of knowledge as of 14 August 2026.

Sources

  • AI Platform Citation Source Index 2026 (5W): consolidation of more than 680 million citations from six studies, August 2024 to April 2026
  • Evertune, analysis of 200 million prompts over five months on citation distribution
  • Tinuiti, AI Citations Trends Q1 2026 (Reddit share by system)
  • Wikipedia, relevance criteria for companies (German-language edition)
  • Wikimedia Foundation, terms of use on disclosure of paid contributions
Free tool

Can ChatGPT find your website at all?

The AI visibility check tests in under a minute whether AI assistants can fetch, understand and cite your page. Eight checks, a concrete fix per finding. No signup.

Run the free check
Matching ArkeonTech service

AI chatbot for sales & support

Answers customer enquiries in seconds, qualifies leads and hands over to your team - live in 2-4 weeks.