Building in Public: Our AI-Citation Tracker Reported Zeros for Clients ChatGPT Was Actually Citing

On July 30 our AI-citation tracker reported that ChatGPT cites trentonsandler.com zero times. It reported the same zero for billybatt.com and matthewjanuszek.com. We asked ChatGPT about those people the same day. It cited Trenton’s own domain as the sole source in its answer. It did the same for Matthew.

The tool was not broken. It was answering a different question than the one we thought we were asking, and we had been reading its answer wrong for weeks. Per our building in public principle, here is what we found, what we changed, and the correction we owe anyone who copied our earlier process.

What we were measuring

Every Wednesday a Claude agent runs a GEO audit across our tracked roster — nine people and companies, including Trenton Sandler, Michael Krigsman of CXOTalk, Western Trading Post, and Igor Ivitskiy. GEO is the answer-engine counterpart to SEO. The question is not “where do you rank” but “when ChatGPT or Perplexity answers a question about you, does it cite your site or somebody else’s?”

We ran that on Ahrefs, which maintains an index of AI citations broken out by engine — ChatGPT, Google AI Overviews, AI Mode, Gemini, Perplexity, Copilot, Grok. It is a good index. It is still an index: a crawled, cached record of citations Ahrefs has observed, not the model’s behavior at the moment a buyer asks.

For a name with a large public footprint the two line up. For a personal brand — the entire category we work in — they diverge, and the index reads low. A zero in an index of AI citations means “not in this corpus.” We were reading it as “invisible to AI.” Those are different claims and only one of them was true.

The second failure, eight days later

On August 9 the weekly run returned no Ahrefs data at all. The workspace had crossed its API unit cap — 104,072 units used against a 100,000 limit — and every call came back API units limit reached. Nothing was misconfigured. A metering threshold reset in ten days, and a report that depended on it had nothing to say.

One weak signal is a measurement problem. A weak signal that can also be switched off by a billing meter is a design problem.

What we switched to

We moved the primary instrument to DataForSEO’s AI Optimization API, which sends a prompt to the live model and returns the answer with its citations attached. We ask ChatGPT and Perplexity directly, then read the annotations array on the response — every cited URL, in order. Whether the client’s own domain got cited becomes a mechanical check instead of a judgment call. The fan_out_queries field shows the retrieval path the model took to get there, which is how you diagnose why an answer went wrong rather than just noting that it did.

Ahrefs stays in the run for exactly one job: Grok and Copilot, which DataForSEO does not cover, and which matter — Grok has been the single largest citation source for CXOTalk. It is now an optional enrichment at the end of the run. If it is over quota, the run notes the reset date and carries on undegraded.

The cost comparison is not close. The full nine-entity roster on live probes runs about $0.43. The same roster on Ahrefs costs roughly 1,080 API units for the weaker signal. We considered upgrading the plan to restore Brand Radar’s per-engine breakdown. Measuring better turned out to be cheaper than measuring more.

Four things the new method found that the old one could not

Identity is not discovery, and only one of them is revenue. Ask ChatGPT “what is Western Trading Post?” and it cites westerntradingpost.com as the sole commercial source. Clean win. Ask it where to buy authentic vintage Native American turquoise jewelry at auction — a question a real buyer asks, with no brand name in it — and the company has zero presence. Not cited, not mentioned, not named in any ranked recommendation. A citation count aggregates those two states into one number and hides the gap that costs money. We now run both prompts for every entity, every week.

A competitor list nobody has tested against an engine is fiction. Our roster had tracked six competitors for that account. When we finally put the discovery question to a live model, none of the six appeared. The names that came back were Santa Fe Art Auction, Morphy Auctions, Bonhams, Christie’s and Turquoise Village. We had been measuring share of voice in a race the engine was not running.

Clients compete against their own duplicate properties. Ask about Igor Ivitskiy and ChatGPT pulls his credentials — the patents, the publications, the 2018 President of Ukraine prize — from ivitskiy.github.io, a GitHub Pages site, and gives his actual domain a single closing citation. A count would show ivitskiy.com being cited. It would not show that a free subdomain is beating it on his own name. The fix is a canonical tag and an afternoon of work, and we would never have looked for it.

Engines disagree about the present, and that disagreement is the sales asset. Trenton Sandler transferred from LSU to Arkansas. In the same week, ChatGPT still described him as an LSU athlete while citing his university roster page. Perplexity had the transfer right — and cited our article about it as the source. Same person, same week, two engines, one of them current because it reads what we publish. No aggregate number produces that sentence.

The correction we owe

Our task file told the agent to label its numbers “Brand Radar per-engine breakdown unavailable on current plan.” That was wrong. The Site Explorer endpoint returns a full per-engine split on the plan we are on. Several weekly digests carried a disclaimer describing a limitation we did not have.

The deeper failure is more useful. That instruction sat in the task file from July 17. We discovered it was wrong on July 30, wrote it down in the agent’s memory, and did not go back and fix the file. So the August 9 run read a stale instruction, rediscovered the same correction from memory, and would have done it again every week. A correction that lives only in a run log is not a correction. It has to go back into the document the work runs from, or you pay for the same discovery on a schedule.

The checklist (steal this)

  1. Probe the live model, not only a citation index. Indexes undercount personal brands, which is the whole category.
  2. Run two prompts per entity: an identity question with the name in it, and a discovery question a buyer would ask with the name left out.
  3. Read the citations off the response’s annotations array. Own-domain-cited is a mechanical check, not an impression.
  4. When the domain is absent, record which domains took the slots. Separate the self-inflicted causes — a duplicate property, the company site outranking the person — from genuine third parties.
  5. Check the answer’s factual claims against the client’s live site. A stale fact in an AI answer is a content assignment.
  6. Query at least two engines. Where they disagree, you have found both a problem and the proof that publishing fixes it.
  7. Pin the model version. Do not upgrade to a newer model mid-series just because one exists — you lose week-over-week comparability.
  8. Where the index and the live probe disagree, the live probe wins, and the report says so.
  9. When a run corrects the process, edit the task file in that run.

Why we publish this

This is LDT — Learn, Do, Teach — and CCS, Content, Checklist, Software. We learned by running the audit weekly, we are teaching by publishing the method with the numbers that forced the change, and the checklist above is already encoded in the agent that runs every Wednesday morning. Publishing the version where our own tooling misread its own data is the part that makes the rest credible.

What we are improving next

Gemini and Claude are available through the same endpoint and are not in the weekly run yet. The roster covers nine entities against a portfolio well past sixty, so it is being extended. And the discovery prompts need to be written and pinned per entity rather than composed at run time, so the week-over-week comparison holds. When those ship, this article gets updated — that is the deal we make by building in public.

Want to see where you stand in AI answers? Get a free Quick Audit.

Dennis Yu
Dennis Yu
Dennis Yu is the CEO of Local Service Spotlight, a platform that amplifies the reputations of contractors and local service businesses using the Content Factory process. He is a former search engine engineer who has spent a billion dollars on Google and Facebook ads for Nike, Quiznos, Ashley Furniture, Red Bull, State Farm, and other brands. Dennis has achieved 25% of his goal of creating a million digital marketing jobs by partnering with universities, professional organizations, and agencies. Through Local Service Spotlight, he teaches the Dollar a Day strategy and Content Factory training to help local service businesses enhance their existing local reputation and make the phone ring. Dennis coaches young adult agency owners serving plumbers, AC technicians, landscapers, roofers, electricians, and believes there should be a standard in measuring local marketing efforts, much like doctors and plumbers must be certified. He has appeared on 353 podcasts with 619 credited episodes — see the full list of his podcast appearances.