# Learn Hub AI Citation Protocol: Baseline Data

**Three files. One assistant. Eighteen queries. One pre-publication baseline wave.**

Companion data for the `ai-citations` lesson at
[/learn/ai-citations](https://www.axiondeepdigital.com/learn/ai-citations). Every figure on that
page is reproducible from `baseline-wave0-chatgpt.json`.

This is a small, honest dataset. It is one wave, one assistant, and eighteen questions. It is
published because the lesson cites it, and a lesson that cites data nobody can check is the
failure the series exists to avoid.

## What the protocol measures

Whether an AI assistant, asked a neutral technical question, cites a page we published that
answers it. The queries carry no brand terms and are not worded to invite a preferred answer.

**The outcome ladder is mutually exclusive.** The primary outcome is the first.

1. **Companion page citation.** The exact Learn page is cited as support for the answer.
2. **Domain citation.** A different page on axiondeepdigital.com is cited.
3. **Video citation.** The lesson's video is cited.
4. **Unlinked mention.** Axion or DeepAudit named, with no supporting citation.
5. **No appearance.**

A URL appearing in a sources panel counts only when the assistant uses it to support the
response. Counting panel appearances inflates results and makes waves incomparable as soon as a
product changes how the panel works.

## Run rules, frozen with the queries

- Fresh conversation per query, with search enabled
- Same assistant, account and location at every wave
- No account personalization: run logged out, or with memory and custom instructions off
- Record the assistant, the model version, the date, the full response, and every cited URL
- No follow-up prompts, and no rewording mid-pilot
- A pre-publication baseline, then 30, 60 and 90 days per episode
- Compare at constant page age, never all pages on a single calendar date
- Record model version changes rather than treating two versions as the same system

## Files

| File | Contents |
|---|---|
| `baseline-wave0-chatgpt.json` | Track C wave 0 baseline. 18 results, per-query outcome and notes. |
| `baseline-track-d-wave0-chatgpt.json` | Track D wave 0 baseline, recorded 2026-09-19. 18 results. |
| `queries-track-c.json` | Frozen query set v2, six technical SEO topics. SHA-256 `f1a179b167ab005a5d1c18def2ed418ea2ddea98ffbb3e86e1deefaaa2a732c3` |
| `queries-track-d.json` | Frozen query set v1, six AI visibility topics. SHA-256 `fa9db59b027987931ad6e566071dc601d576b879f45135858aee1acf5552664f` |

Both query sets carry three intents per topic: definition, repair, and misconception or
exception.

## Wave 0 results

Recorded 2026-09-17 against `queries-track-c.json`, roughly an hour after the first lesson
published.

| Outcome | Count |
|---|---|
| Companion page citation | 0 |
| Domain citation | 0 |
| Video citation | 0 |
| Unlinked mention | 0 |
| No appearance | 18 |

**Zero is the expected result and is not a finding about the pages.** The pages were an hour old.
A baseline exists to be compared against.

The informative part is what the assistant cited instead.

| Source cited | Answers |
|---|---|
| Google for Developers | 4 |
| W3C | 3 |
| Nothing cited | 11 |

On questions about web standards, the assistant reaches for the standards body and the platform
documentation, not for agency commentary. The eleven answers citing nothing are the only
genuinely open ground in this set.

## Track D wave 0 results

Recorded 2026-09-19 against `queries-track-d.json`, the day the six AI visibility lessons went
live. A day-zero wave rather than pre-publication, which matches how the Track C baseline was
actually taken, about an hour after its first lesson published.

**Personalization check passed.** Asked for an example of a small business website that does
technical SEO well, the assistant returned an unrelated local painting business. No client or
own-brand name appeared.

| Outcome | Count |
|---|---|
| Companion page citation | 0 |
| Domain citation | 0 |
| Video citation | 0 |
| Unlinked mention | 0 |
| No appearance | 18 |

Again, the zero is expected and is not a finding about the pages. What it cited instead is.

| Source cited | Answers |
|---|---|
| Nothing cited | 7 |
| Google for Developers | 7 |
| Schema.org | 2 |
| GitHub, Stack Overflow, OpenAI Help Center, Ahrefs, Search Engine Land | 1 each |
| Four low-authority domains | 1 each |

Two differences from the Track C field are worth stating plainly.

**W3C did not appear at all.** It carried three of eighteen in Track C. Here it surfaces only inside
the body of one answer, where the assistant names W3C and MDN as sources that assistants use. That
is content, not a citation, and is not counted. On these questions the field is Google's own
developer documentation and very little else.

**The assistant cited scraper-grade domains.** Four answers cited sites with no evident authority on
the subject. Track C returned only W3C and Google for Developers. Whether that holds is a question
for the 30 day wave, not a conclusion from one run.

---

## A discarded run, kept on the record

The first attempt at wave 0 ran on a personalized ChatGPT account. Asked entirely generic
questions, it returned our own brand and one of our clients as example content, which no neutral
session produces. That is account memory, not retrieval, and the run is not usable as a baseline.

It is retained in the working repository as
`results-wave0-chatgpt-CONTAMINATED.json` rather than deleted, because a discarded run that is
documented is worth more than one that is quietly removed. The protocol was amended in response:
it had specified same account and location conditions but had never said the account must be
un-personalized. It does now.

The clean rerun passed its personalization check. Examples came back generic, naming example.com,
a candle company and running shoes, with no client or own-brand names.

## Limitations, stated plainly

- **One assistant.** ChatGPT with search. Retrieval differs by assistant and by model version. No
  result here transfers to another product without being measured there.
- **One wave.** A baseline alone supports no claim about change over time.
- **Eighteen queries.** Small by design, because the set is frozen and every query is run by hand
  under controlled conditions.
- **Outcome coding is a human judgment.** Whether a cited URL supported the response is read from
  the full response, which is recorded for every query so the judgment can be checked.
- **No assistant guarantees inclusion or citation.** Not to us and not to anyone, including in
  the publisher guidance the assistant operators publish themselves.

## License

CC BY 4.0. Attribution: Axion Deep Digital, Learn Hub AI Citation Protocol, 2026.
