Vol. 01 — AI Citation IndexAugust 2026
Citation/Research

Who does AI cite when people ask? We measure it — industry by industry.

Research protocol

Methodology

Every figure on this site is produced by the protocol below — published in full, because research you can't inspect is marketing. Version 1.3, effective July 2026.

01The question we answer

When a person asks an AI answer engine a question in a given industry, which sources does the engine cite — and how does that change over time? We measure citation behavior, not answer quality. A citation is any linked or named source attached to an answer, whether it appears inline, in a source list, or in a 'learn more' module.

We publish observational measurements of platform behavior. Nothing on this site is an endorsement, a ranking of quality, or a claim about any company's business — only about what the engines surfaced, on the dates we measured, under the protocol described here.

02Prompt corpus construction

Each industry corpus starts from real query demand: search-console exports, People-Also-Ask trees, autocomplete data, community questions (Reddit, Quora, industry forums) and keyword research tools. Candidate questions are deduplicated, normalized, and classified into intent classes.

For US real estate, 612 prompts span six classes: buying process (21%), selling (14%), financing & mortgages (26%), regulation & rights (9%), market sentiment (18%), and local discovery (12%). Prompts are weighted by estimated US monthly search volume so that citation share reflects visibility on the questions people actually ask — not a flat average over questions nobody asks.

The corpus is frozen between monthly runs. Additions and removals are logged in a public changelog; trend series are always computed on the frozen set.

03Sampling protocol

Every prompt is run 5 times per platform per measurement window (the first full week of each month). Runs use fresh, unauthenticated sessions with no chat history, no personalization, and default settings. Search-enabled configurations are used where the platform offers them.

Repeated runs are not optional. Answer engines are stochastic systems — a single run of a single prompt tells you almost nothing. Five runs per prompt per platform yields 15,300 answer runs per industry per month, which is what allows us to attach confidence intervals to every share estimate we publish.

04Citation extraction

Citations are extracted deterministically from answer text, footnote lists and source modules — no model-assisted judgment is used in the extraction step itself. URLs are normalized and mapped to registered domains; subdomains roll up to the parent domain (e.g. www.zillow.com → zillow.com), with the exception of platform-hosted communities (reddit.com, youtube.com), which are tracked at the domain level by design.

A 4% stratified random sample of every month's extraction output is audited by hand against the raw answers. In the July 2026 audit, extraction precision measured 98.2% and recall 97.6%. Audits are archived and available on request.

05Metric definitions

Citation share (per platform): of all citation events recorded for that platform, the percentage pointing to a given domain. Blended citation share: the usage-weighted mean of per-platform shares, using estimated US AI-search referral traffic per platform as weights. Weights are reviewed quarterly and published with the dataset.

Citation rate: the percentage of answer runs in which a domain was cited at least once. Average first-citation position: the mean ordinal position of a domain's first citation within an answer's source list (1 = cited first). Δ (delta): change in blended citation share versus the previous measurement window, in percentage points.

06Confidence & limitations

At 15,300 runs per industry-month, blended citation-share estimates carry a 95% confidence margin of approximately ±1.2 pp for a domain at 15% share; margins are smaller for smaller shares and are published per-domain in the dataset. Month-over-month changes smaller than the margin should be read as noise.

Known limitations: (1) results reflect US-English, US-locale sessions and do not generalize to other locales; (2) platform weights are estimates and shift the blended figures when revised; (3) engines update models continuously, so intra-month behavior may drift between measurement windows; (4) personalization is disabled, so logged-in user experiences may differ. We publish our raw data precisely so others can test these sensitivities.

07Refresh cadence

Each tracked industry is re-measured monthly, in the first full week of the month, on the frozen prompt corpus. Reports are published within five business days of the measurement window closing. Historical datasets remain available; we never retroactively edit published figures — corrections are issued as errata with a public log.

08Reuse & citation

All datasets are released under CC-BY 4.0. You may reuse, remix and republish them — including commercially — with attribution. Journalists: the fastest way to cite a figure is 'according to the Citation Research AI Citation Index (July 2026)' with a link to the relevant industry page.

Suggested citation: Citation Research Lab (2026). US Real Estate AI Citation Index, July 2026. citationresearch.com/industries/us-real-estate.

See it applied

The protocol above produced the US Real Estate AI Citation Index and the 2026 flagship report. Every number there traces back to this page.

Free visibility scan

Is AI citing your brand — or your competitor?

This index is a monthly snapshot. Pagelens tracks your citation share across every AI engine in real time, down to the prompt.

Powered by Pagelens. Free, no account required.