Skip to content
← Journal

2 September 2026 · 13 min

AU AI Search Watch, Edition 2

Across 13 buying questions in four Canberra trades, the AI engines never once named a verified local business for 'best plumber in Canberra' or its equivalent. 90 questions, seven query sets, six trades, and the first month-over-month comparison this series has been able to make.

Four long rows of small pale paper markers run across a near-white surface, and a single empty column cuts down through all four where one marker is missing from every row, a burnt-orange plumb bob hanging weightless above the gap on a dotted line and casting a faint glow into it: four trades measured separately, failing at exactly the same question.

Measured 1 September 2026. This is the second edition of a monthly series. Edition 1 is here, and it is worth reading first, because this edition exists to be compared against it.

Summary

On 2 September 2026 we ran 90 customer-style questions through Google AI Overview and ChatGPT search across six Australian local-services categories. Four of those categories are new this edition: Canberra electricians, plumbers, roofers and landscapers, each asked the same thirteen questions in the same order so the trades can be read against each other.

The number we did not expect: on the plainest question a customer can ask, "best plumber in Canberra" and its equivalent in each trade, the engines named a verified local business in zero of eight cells. Four trades, two engines, nothing. Meanwhile the same engines named a local business on 9 to 11 of the 13 questions overall.

The two engines still barely cite the same sources, 16.5% domain overlap against Edition 1's 11.7%. This is one operator's snapshot, not a market study.

Two things this edition promised and did not deliver are named in full further down, under What we promised and could not deliver. They are a missing Bing measurement and a missing prose-mention tally, both for the same reason: the instrument does not exist yet. Neither is buried, because the value of this series is that it reports its own limits before anyone else has to find them.

What changed since Edition 1

Edition 1 could not compare anything to anything. It said so plainly: "No number on this page can be compared to a previous number, because there is no previous number." That sentence is now out of date, and this is the first edition where a trend is even arguable.

Edition 1Edition 2
Measured2 August 20261 September 2026
Query sets37
Queries3890
Trade categories3 (cat grooming, house painting, tradie marketing)6 (those three, plus electricians, plumbers, roofers, landscapers)
Runs per query11 on two sets, 3 on five sets
Engine checks attempted76420
SpendUSD $0.28USD $1.49

The 38 original queries are unchanged, word for word. That is the point of freezing them.

Method, stated plainly

ItemDetail
Date of the run1 September 2026
Queries90
Query sets7 (cat grooming Melbourne 8; house painting Canberra 22; SEO and websites for tradies Canberra 8; and four Canberra trade probes of 13 each: electricians, plumbers, roofers, landscapers)
Engines2 (Google AI Overview, ChatGPT search)
Runs per query3 for cat grooming and all four trade probes, 1 for house painting and tradie marketing
Engine checks attempted420
Checks that failed on the provider side10, all on the Google leg, all `DataForSEO 40101: Internal SE Server Error`. Named in full below.
Checks that returned an answer324 of 420. A further 86 returned no AI Overview at all, which is a result and not a failure.
Citations captured2,139 across 547 unique domains
LocationsGoogle AI Overview at city level: Melbourne for the grooming set, Canberra for all six Canberra sets. ChatGPT search is country level only, so every set ran as "Australia".
Data sourceDataForSEO SERP `ai_overview` references for Google, DataForSEO LLM Scraper for ChatGPT search
Total spend on the datasetUSD $1.49

How the queries were phrased

Queries were written the way a customer types or speaks, not as keyword strings. The four new trade sets each reuse the same thirteen slots in the same order, so the trades can be read against each other question by question: slot 1 is the emergency transactional question in every set, slot 8 the regulation question, slot 12 the generic "best X in Canberra".

Examples, verbatim from the frozen sets: "emergency plumber Canberra open now", "my roof is leaking after a storm who do I call Queanbeyan", "do I need approval to build a retaining wall over a metre in the ACT", "who installs EV chargers at home in Canberra".

What changed in the method since Edition 1

  • Three runs per query on five of the seven sets. Edition 1 ran everything once and named that as its weakest property. Majority scoring across three runs is how this edition answers it.
  • Four new trade categories, frozen before they were ever run, so their baselines are triple-run from birth rather than upgraded later.
  • One thing did not change and must not: the original 38 queries, their wording, their order and their location codes.

Sample size limits, said up front

  • 90 queries is still a small sample. It is six trade categories in two cities. It is not representative of Australian local search and should not be cited as if it were.
  • Two measurement rounds is not a trend. It is two points. A direction is not established by two readings, and any movement below is reported as movement between two dates, not as a trajectory.
  • The four trade sets have one reading each. They were frozen for this edition, so they have a baseline and nothing to compare it to. Their numbers describe September and no other month.
  • Three of the 90 queries' subjects are businesses tyso works with (ANT Painting is a family business Tyson co-runs; tyso.io is our own site; Sophisticated Manes is a client). Their domains are marked wherever they appear.
  • This is a snapshot of what engines cited, not of what customers ask. No AI query-volume data exists for Australia, or anywhere.
  • API answers are not app answers. Personalisation, memory and app-side retrieval mean chatgpt.com can answer differently from the API path measured here.
  • Ten calls failed and are excluded from every rate rather than counted as zeroes. All ten were Google AI Overview checks returning `DataForSEO 40101: Internal SE Server Error`: cat grooming cost Melbourne (run 1) and best cat groomer south east Melbourne (run 3); EV charger cost ACT (run 1); leaking hot water Belconnen and blocked drain Gungahlin (both run 3); metal roofing cost ACT (runs 1 and 2); backyard that floods Queanbeyan (runs 1 and 3); turf cost Canberra (run 2). ChatGPT search failed 0 of 210.
  • A first attempt on 1 September was discarded, and it is worth saying why. That run hit a provider outage on the ChatGPT leg which failed between 62% and 85% of its cells while the Google leg failed almost none. Not one electrician cell had three usable ChatGPT runs. Rather than publish a number resting on the surviving fragments, we re-ran the whole panel the next day. The failed attempt is kept with the raw files.

Source classification

Domains were sorted by hand into directory or marketplace, forum or social or video, business website, cost guide, and government or industry body. The classification is a judgement call, so the domains in each bucket are named and the raw citation lists are in the source files. Anyone can re-sort them.

For the four trade probes a second hand pass was needed, and it is the slowest part of this edition: deciding whether each cited business is genuinely a Canberra-region operator. Directories, government, energy retailers, media and cost-guide content never count. A business that cannot be placed is unverifiable and is never assumed local. Every named verdict rests only on verified-local businesses, and the ledgers are published with the run files.

Finding 1: what moved between August and September

ANT Painting was named with a link on 9 of its 22 questions, against 7 on 2 August. That plus two is the least interesting way to say what happened.

SetSurface2 August2 September
ANT PaintingGoogle AI Overview7 of 225 of 22
ANT PaintingChatGPT search0 of 217 of 22
tyso.ioGoogle AI Overview0 of 80 of 8
tyso.ioChatGPT search0 of 80 of 8
Sophisticated ManesGoogle AI Overview0 of 70 of 8
Sophisticated ManesChatGPT search0 of 83 of 8

Underneath the union figure, ANT's two surfaces moved in opposite directions. ChatGPT went from citing the domain zero times to citing it seven. Google AI Overview went the other way, seven down to five. A single "AI visibility" number would have reported mild growth and hidden both halves of what actually happened.

Sophisticated Manes gained three ChatGPT citations somewhere between early and late August and has held exactly 3 of 8 across three separate measurement dates since, including a majority of three runs here. Stability across dates is a stronger statement than the number itself.

tyso.io, our own site, is 0 of 8 cited and 0 of 8 mentioned on both engines, unchanged from August. We publish that because a measurement series that hides its own worst number is not worth reading.

Finding 2: does the engine-overlap gap hold?

Edition 1's central number was that Google AI Overview and ChatGPT search shared only 11.7% of the domains they cited, and the practical reading was that "AI visibility" as a single score hides more than it shows.

Edition 1 found 11.7% overlap and read it as evidence that "AI visibility" cannot honestly be sold as one score. This run pools 16.5%. The gap narrowed, and the finding holds: five sixths of what one engine cites, the other does not.

Query setAIO domainsChatGPT domainsSharedOverlap
tyso.io453811.2%
Cat grooming, Melbourne3629610.2%
House painting, Canberra64591716.0%
Electricians, Canberra67721512.1%
Plumbers, Canberra42902422.2%
Roofers, Canberra36681820.9%
Landscapers, Canberra61742320.5%
All seven pooled2993389016.5%

The tyso.io row is the outlier and it is our own set: 45 domains from Google, 38 from ChatGPT, exactly one in common. Whatever these engines think the "SEO for tradies" market is, they do not think the same thing. | | | | |

Finding 3: six trades, read slot by slot

This is what the four new sets were built for. Because every trade set uses the same thirteen slots in the same order, the same question can be read across trades instead of only within one.

Y means a verified local business in that trade was named on a majority of three runs. A dot means it was not. A question mark means the cell was excluded after a provider failure. G is Google AI Overview, C is ChatGPT search.

#The question's jobElectriciansPlumbersRoofersLandscapers
G / CG / CG / CG / C
1Emergency, transactional. / Y. / YY / Y. / .
2Symptom-first emergency, QueanbeyanY / .Y / YY / .? / .
3After hours, Tuggeranong. / Y. / Y. / Y. / Y
4Big job, cost framingY / Y. / .. / YY / .
5Big job, hire intent, Belconnen. / Y. / Y. / Y. / Y
6Mid job, cost framing. / YY / YY / YY / Y
7Mid job, hire intent, Gungahlin. / Y. / Y. / YY / .
8Regulation question, ACTY / .Y / .Y / .Y / .
9Regulated job, hire intent, Woden. / .. / YY / YY / Y
10The growth jobY / YY / YY / Y. / .
11The growth job at cost stageY / .Y / .? / YY / Y
12"Best X in Canberra". / .. / .. / .. / .
13Near me, licensed. / Y. / .. / .. / Y

Question level: electricians 11 of 13, roofers 11 of 13, plumbers 10 of 13, landscapers 9 of 13.

Three things fall out of reading down the columns rather than across the totals.

Slot 12 is empty everywhere. "Best electrician in Canberra", "best plumber in Canberra", "best roofer in Canberra", "best landscaper in Canberra". Four trades, two engines, eight cells, and a verified local business was named in none of them. This is the question a business owner is most likely to type when checking on themselves, and it is the one question in the set where the engines reliably answer with directories, review platforms and listicles instead of a business. Whatever else is true, being good at this question is not what the AI layer currently rewards.

Slot 8 is the mirror image, and it is unanimous the other way. On the regulation question every trade scored Y on Google and a dot on ChatGPT. Four for four, both directions. Google's AI Overview reaches for a local operator on "do I need approval to build a retaining wall over a metre in the ACT" and its siblings; ChatGPT reaches for the regulator.

Landscaping broke where it was predicted to break. This set was frozen with a stated expectation: it has no genuine panic-buy in slots 1 to 3, so if the citation pattern depends on emergencies, landscaping is where it should fail. It did, and only there.

TradeEmergency slots 1 to 3Everything else, slots 4 to 13
Electricians3 of 610 of 20
Plumbers4 of 69 of 20
Roofers4 of 611 of 19
Landscapers1 of 511 of 20

Landscaping is last on the emergency slots by a distance and first-equal on everything else. It is not a weaker category. It is a category without an emergency, and the emergency questions are where local businesses get named.

Finding 4: how much of this is noise?

Edition 1 ran each query once and said a single run is a snapshot, not a distribution. Five of the seven sets ran three times here, so run-to-run churn is measurable for the first time.

Five of the seven sets ran three times. Across those five, 102 of 118 usable cells returned the same verdict in all three runs.

SetIdentical in all three runsRuns disagreed
Cat grooming, Melbourne16 of 160
Electricians, Canberra25 of 261
Plumbers, Canberra22 of 264
Landscapers, Canberra21 of 254
Roofers, Canberra18 of 257

The grooming set did not waver once, which is why we are willing to say in a client report that its three ChatGPT citations are real rather than lucky. Roofing is the least stable of the six trades, disagreeing on 7 of 25 cells, and any single-run reading of a roofing question should be treated accordingly.

This is the answer to Edition 1's own criticism of itself. One run per query was the weakest thing about that edition. On this evidence most cells are stable, but 16 of 118, about one in seven, are not, and the unstable ones cannot be told apart from the rest without running the query more than once.

Finding 5: did AI Overviews show up for find-a-provider questions?

Edition 1 found that cost, comparison and licensing queries returned AI Overviews while find-a-provider queries mostly did not, and read that as the AI layer sitting upstream of the hiring decision rather than on top of it.

It held, and more sharply than in Edition 1.

86 of 210 Google AI Overview checks, 41%, returned no AI Overview at all. ChatGPT search answered 210 of 210. When Google's AI layer declines to appear, the ordinary results page is what the customer sees, and the question of who the AI cites does not arise.

The trade sets make the pattern legible because each pairs a cost-framed question with a hire-intent question on the same job. Slots 4, 6 and 11 are cost questions; slots 5, 7 and 9 are the same jobs at hire intent. Google names a local business more readily on the cost side, ChatGPT more readily on the hire side, and slot 5, the big job at hire intent in Belconnen, is a dot on Google in all four trades while ChatGPT scores it Y in all four.

Edition 1 read this as the AI layer sitting upstream of the hiring decision rather than on top of it. Two engines now look less like one layer and more like a division of labour: Google answering the research question, ChatGPT answering the who-do-I-call one.

Cited domains, counts

Every domain cited on 8 or more of the 90 queries. "Queries" counts distinct queries, so a domain cited twice in one answer counts once.

DomainTypeQueries (of 90)Google AIO citationsChatGPT citations
hipages.com.auMarketplace25922
whatsthedamage.com.auCost guide19129
reddit.comForum18173
thequoteyard.com.auMarketplace17613
youtube.comVideo15150
facebook.comSocial14115
region.com.auLocal media1378
airtasker.comMarketplace1249
starworks.com.auReview platform12112
planning.act.gov.auGovernment12103
google.comMap entity cards11110
yellowpages.com.auDirectory1139
wordofmouth.com.auDirectory11210
localsearch.com.auDirectory1028

Not one business website appears in that list. Every domain cited on 8 or more of the 90 questions is a marketplace, a directory, a forum, a video platform, a cost guide, a news masthead or a government page. Edition 1 found directories dominant on ChatGPT and forums and video dominant on Google; both halves replicate here, and the split is visible in the last two columns. | | | | |

Domains marked * are businesses tyso works with.

The trade probes: how often did AI name a local business?

Each figure below is majority-scored across three runs, and every business behind it was checked against its own website before it counted. Businesses in the right region but the wrong trade do not count. Businesses whose region we could confirm but whose trade their own site never states are recorded as unverifiable and are not counted as local: that was 35 domains across the four sets, and counting them would have inflated every number here.

Electricians. Across 13 buying questions they named a local sparkie 11 times. Google AI Overview 5 of 13, ChatGPT search 8 of 13. In August the same set scored 7 of 13, but August excluded four cells to this month's none, so the denominators are not the same and part of that rise is a better measurement rather than a better month.

Roofers. Across 13 buying questions they named a local roofer 11 times. Google AI Overview 6 of 12, ChatGPT search 9 of 13. The highest ChatGPT score of the four trades, and also the least stable set in the panel.

Plumbers. Across 13 buying questions they named a local plumber 10 times. Google AI Overview 5 of 13, ChatGPT search 8 of 13.

Landscapers. Across 13 buying questions they named a local landscaper 9 times. Google AI Overview 6 of 12, ChatGPT search 6 of 13. The lowest of the four, and as the slot table shows, the shortfall is entirely in the three emergency-shaped questions the trade does not really have.

What this suggests for Australian local businesses

Stated as observation from one small sample, not as advice that has been proven to work.

  1. The question you would check first is the one question nobody wins. "Best [your trade] in Canberra" named a verified local business in zero of eight cells. If you are judging your own AI visibility by typing that question, you are judging it by the least informative question in the set.
  2. Being cited is not one thing, and it is still not one thing. 16.5% domain overlap between the two engines. Anyone selling a single AI visibility score should be asked which engine it measures, and what it says about the other.
  3. The emergency is where local businesses get named. Across the three trades that have genuine emergencies, 11 of 18 emergency cells named a local business. Landscaping, which has none, managed 1 of 5 while matching everyone else on the other ten questions.
  4. The two engines appear to divide the labour. Google's AI Overview leans to the research and regulation questions and did not appear at all on 41% of checks. ChatGPT answered every time and leans to the who-do-I-call questions.
  5. A surface can arrive from nothing. ANT Painting went from 0 to 7 ChatGPT citations in a month while losing two on Google. If it had been tracked as one merged number, the arrival of an entire surface would have shown up as a rounding error.
  6. Directories still hold the ground. Every domain cited on 8 or more of the 90 questions is a marketplace, directory, forum, video platform, cost guide, news site or government page. Not one is a business's own website.
  7. We measured citations, not customers. Nothing on this page says any of it produced a phone call, and we make no revenue claim of any kind.

What we promised for this edition, and what we could not deliver

Edition 1 published five commitments so they could be checked. Three were delivered, two were not, and the two failures are worth more to a reader than the three successes.

#Committed in Edition 1Outcome
1The same 38 queries, same locations, same engines, re-runDelivered. The panel grew to seven sets and 90 queries, and the original 38 are untouched.
2Three runs per query, majority-scored, on at least one setDelivered on five sets.
3At least one new trade categoryDelivered four: electricians, plumbers, roofers, landscapers.
4Bing, retriedNot delivered.
5A count of how often each engine names a business without citing itNot delivered.

On Bing. Edition 1 promised a retry because ChatGPT retrieval draws on Bing's index, and the 3 August attempt had returned mismatched results. The retry did not happen, and the honest reason is that the measurement tool has never had a Bing runner in it. It supports two engines, Google AI Overview and ChatGPT search, and that is all it has ever supported. The commitment was not keepable as written at the time it was written. Bing stays on the list for Edition 3, and it will only go back on that list once the runner exists rather than once the intention does.

On the prose-mention tally. Edition 1 noticed that engines sometimes name a business without linking to it, and promised to count it. The field that would carry that count comes back empty on every ChatGPT cell of every set, across four separate measurement dates. This is not one run missing it. The instrument reports citations and cannot currently see a name that carries no link, so the count does not exist and no number is offered in its place. Being named without being linked remains, in our view, the more interesting number of the two, which is why it is worth saying loudly that we still cannot measure it.

Neither of these is presented as a small thing. A measurement series that quietly drops the commitments it cannot meet is worth less than one that keeps a public tally of its own failures.

What would change the picture

  • Two rounds is not a trend, and four of the seven sets have only one round. The trade probes were frozen for this edition. Their numbers describe September and nothing else.
  • The provider is part of the instrument. The first attempt at this edition was thrown away because one engine's leg failed on most of its cells. A quieter version of the same problem, failing 10% rather than 70%, would be harder to spot and would bias a month.
  • 41% of Google checks showed no AI Overview. That share moving is enough on its own to change most numbers on this page.
  • The trade sets are Canberra-only and one city cannot generalise. The grooming set is Melbourne, and it behaves differently from everything else here.
  • Two measurements we still cannot make, both described in the section above: no Bing, and no count of businesses named without a link.
  • Hand-scoring is a judgement. 245 domains were checked by a person against each business's own site for this edition, with 22 more carried from August's electrician ledger. The evidence lines are published with the run files so anyone can disagree with a specific call rather than the total.

What Edition 3 will measure

Committed now, so it can be checked later.

  1. The same 90 questions, same locations, same engines, re-run. This is the only commitment that costs nothing but discipline, and it is the one that makes the series worth reading.
  2. The four trade sets get their second reading, which turns four birth baselines into four comparisons.
  3. A stated failure rate per engine on every future edition, in the method table, whether or not it is embarrassing. The 1 September attempt would have been much harder to catch without it.

That is the whole list, and it is deliberately shorter than Edition 1's. Two of Edition 1's five commitments have now gone unmet in a row. Bing goes back on this list when the measurement tool has a Bing runner in it, and not before.

Edition 3 is scheduled for early October 2026.

Verification

Every figure on this page comes from a dated file in our measurement archive. The query sets are frozen, one file per set, one query per line, and every citation list is stored raw, so the whole page can be recomputed from the files. The four trade-probe locality ledgers are published alongside the run files, one line of evidence per business. We hand the raw files over on request: tys@tyso.io.

Who made this. tyso is a one-person growth studio in Canberra. This page is one small operator's measurement, run on its own clients, its own site and four Canberra trade categories, published because no Australian equivalent exists. It is not a market study and should not be cited as one.

Corrections. If a number here is wrong, tell us and we will fix it and date the fix. tys@tyso.io

AI searchmeasurementdata